Skip to content

Ch 16: str字符串切片

str是Rust中最基本的字符串类型,代表UTF-8编码的字符串切片。它总是以引用的形式存在(&str),是字符串数据的”视图”,不拥有所有权。

// &str 是指向字符串数据的引用
let s: &str = "hello world";

理解str和String的关系:

  • String - 拥有所有权的可变字符串
  • &str - 不可变的字符串切片引用
fn main() {
let s = "apple,banana,cherry";
// split 按分隔符分割,返回迭代器
let parts: Vec<&str> = s.split(',').collect();
println!("{:?}", parts);
// 输出: ["apple", "banana", "cherry"]
// splitn 限制分割次数
let parts: Vec<&str> = s.splitn(2, ',').collect();
println!("{:?}", parts);
// 输出: ["apple", "banana,cherry"]
// rsplitn 从右边开始分割
let parts: Vec<&str> = s.rsplitn(2, ',').collect();
println!("{:?}", parts);
// 输出: ["cherry", "apple,banana"]
}
fn main() {
let s = "hello123world456rust";
// 按字符类别分割
let parts: Vec<&str> = s.split(|c: char| c.is_numeric()).collect();
println!("{:?}", parts);
// 输出: ["hello", "world", "rust"]
// split_once 分割一次,返回Option
let s = "key=value";
if let Some((key, value)) = s.split_once('=') {
println!("key: {}, value: {}", key, value);
}
// rsplit_once 从右边分割一次
let path = "/home/user/file.txt";
if let Some((dir, file)) = path.rsplit_once('/') {
println!("dir: {}, file: {}", dir, file);
}
}
fn main() {
let multiline = "line one\nline two\r\nline three";
// lines 按行分割(识别\n和\r\n)
for (i, line) in multiline.lines().enumerate() {
println!("行{}: '{}'", i + 1, line);
}
// 输出:
// 行1: 'line one'
// 行2: 'line two'
// 行3: 'line three'
// lines返回&str,可以继续调用其他方法
let first_line = multiline.lines().next().unwrap_or("");
println!("第一行: {}", first_line);
}
fn main() {
let s = "hello";
// chars 返回字符迭代器
for c in s.chars() {
print!("{} ", c);
}
println!();
// 输出: h e l l o
// char_indices 返回(字节位置, 字符)
for (i, c) in s.char_indices() {
println!("'{}' at byte {}", c, i);
}
// 输出:
// 'h' at byte 0
// 'e' at byte 1
// 'l' at byte 2
// 'l' at byte 3
// 'o' at byte 4
}

重要:UTF-8中汉字通常占3个字节:

fn main() {
let chinese = "你好";
println!("len (bytes): {}", chinese.len()); // 6
println!("chars count: {}", chinese.chars().count()); // 2
// 遍历字符和字节位置
for (i, c) in chinese.char_indices() {
println!("'{}' at byte {}", c, i);
}
// 输出:
// '你' at byte 0
// '好' at byte 3
}
fn main() {
let s = "hello";
// bytes 返回字节迭代器
for b in s.bytes() {
print!("{} ", b);
}
println!();
// 输出: 104 101 108 108 111
// 汉字的字节
let chinese = "你好";
println!("bytes: {:?}", chinese.bytes().collect::<Vec<_>>());
// 输出: [228, 189, 160, 229, 165, 189]
}
fn main() {
let s = " hello world ";
// trim 去除两端空白
println!("trim: '{}'", s.trim());
// 输出: 'hello world'
// trim_start / trim_end 只去除一端
println!("trim_start: '{}'", s.trim_start());
println!("trim_end: '{}'", s.trim_end());
// trim_matches 去除指定字符
let s2 = "!!!hello!!!";
println!("trim_matches: '{}'", s2.trim_matches('!'));
// 输出: 'hello'
}
fn main() {
let s = "The quick brown fox";
// contains 检查是否包含子串
println!("contains 'quick': {}", s.contains("quick"));
// starts_with / ends_with
println!("starts with 'The': {}", s.starts_with("The"));
println!("ends with 'fox': {}", s.ends_with("fox"));
// find 返回第一个匹配的位置
println!("find 'brown': {:?}", s.find("brown"));
// rfind 从右边查找
println!("rfind 'o': {:?}", s.rfind("o"));
}
fn main() {
let s = "hello world";
// get 获取子串(按字节索引)
if let Some(sub) = s.get(0..5) {
println!("get(0..5): '{}'", sub);
}
// 切片的便捷语法
println!("切片[0..5]: '{}'", &s[0..5]);
// split_at 按字节位置分割
let (left, right) = s.split_at(6);
println!("split_at(6): '{}' 和 '{}'", left, right);
}
fn main() {
// String -> &str
let s = String::from("hello");
let s_slice: &str = &s;
let s_slice2: &str = s.as_str();
let s_slice3: &str = &s[..]; // 切片语法
// &str -> String
let s1: &str = "world";
let s2: String = s1.to_string();
let s3: String = String::from(s1);
let s4: String = s1.to_owned();
}
操作复杂度
len()O(1)
chars().count()O(n)
split()O(n)
find()O(n)
get(0..k)O(1)

今天我们学习了&str字符串切片的核心方法:

  1. split系列 - split、splitn、rsplitn、split_once按分隔符分割
  2. lines - 按行分割,识别\n和\r\n
  3. chars - 字符迭代,处理Unicode
  4. bytes - 字节迭代
  5. trim - 去空白操作

理解UTF-8编码和字符/字节的区别是正确处理Rust字符串的关键。