Skip to content

Day 14: 字符串

Rust中的字符串处理是新手最容易困惑的部分。关键在于理解两种字符串类型:

  • String:一个可增长的、堆分配的 UTF-8 字符串类型
  • &str:一个指向 UTF-8 字符串的不可变引用(也称为”字符串切片”)
fn main() {
// String 是拥有的(owned)字符串
let s1: String = String::from("hello");
// &str 是借用的字符串引用
let s2: &str = "hello";
// String 可以被修改
let mut s3 = String::from("hello");
s3.push_str(", world!");
println!("{}", s3);
// &str 是只读的
let s4 = "hello";
// s4.push_str(", world!"); // 编译错误!
// String 可以转换为 &str
let s5: &str = &s1;
println!("s5 from s1: {}", s5);
}

理解Rust字符串的内部结构有助于深入掌握:

fn main() {
// String 的结构类似这样:
// struct String {
// ptr: *const u8, // 指向堆内存的指针
// len: usize, // 字节长度
// capacity: usize, // 容量
// }
// &str 的结构类似这样:
// struct &str {
// ptr: *const u8, // 指向字符串的指针
// len: usize, // 字节长度
// }
let s = "你好";
println!("字符串: {}", s);
println!("字节长度: {}", s.len()); // 6 bytes (每个汉字3字节)
println!("字符数: {}", s.chars().count()); // 2 个字符
}
fn main() {
// 多种创建方式
let s1 = String::new(); // 空字符串
let s2 = String::from("hello"); // 从字面量创建
let s3 = "hello".to_string(); // 使用 to_string()
// 使用 with_capacity 预分配
let mut s4 = String::with_capacity(100);
s4.push_str("hello");
println!("s4 capacity: {}, len: {}", s4.capacity(), s4.len());
// 从其他类型转换
let num = 42;
let s5 = num.to_string();
let s6 = 3.14.to_string();
println!("数字转字符串: {} {}", s5, s6);
}
fn main() {
let s = String::from("hello world");
// 长度和判空
println!("len: {}, is_empty: {}", s.len(), s.is_empty());
// 追加和拼接
let mut s1 = String::from("hello");
s1.push(' '); // 追加单个字符
s1.push_str("world"); // 追加字符串切片
println!("s1: {}", s1);
// 使用 + 运算符(需要 &str,会获取所有权)
let s2 = String::from("hello ");
let s3 = String::from("world");
let s4 = s2 + &s3; // s2 被移动,无法再使用
println!("s4: {}", s4);
// println!("s2: {}", s2); // 编译错误!
// 使用 format! 宏(不获取所有权)
let s5 = String::from("tic");
let s6 = String::from("tac");
let s7 = String::from("toe");
let s8 = format!("{}-{}-{}", s5, s6, s7);
println!("s8: {}", s8);
// 字符串索引(需要谨慎)
let s9 = String::from("hello");
// let c = s9[0]; // 编译错误!Rust字符串按字节索引
let c = &s9[0..1]; // 正确:按字节切片
println!("第一个字符: {}", c);
}

字符串切片是对 String 或字面量的引用,必须在有效的 UTF-8 边界上切分:

fn main() {
let s = String::from("hello world");
// 字符串切片
let hello = &s[0..5];
let world = &s[6..11];
println!("{} {}", hello, world);
// 使用 chars() 按字符遍历
println!("按字符遍历:");
for c in s.chars() {
print!("{} ", c);
}
println!();
// 使用 bytes() 按字节遍历
println!("按字节遍历:");
for b in s.bytes() {
print!("{} ", b);
}
println!();
// 字符串方法
let s2 = String::from(" hello ");
println!("trimmed: '{}'", s2.trim());
println!("uppercase: {}", s.to_uppercase());
println!("lowercase: {}", s.to_lowercase());
println!("contains 'world': {}", s.contains("world"));
println!("starts_with 'hello': {}", s.starts_with("hello"));
println!("replace: {}", s.replace("world", "rust"));
}

Rust的 String 类型保证是有效的 UTF-8。但在处理其他编码时需要额外注意:

fn main() {
// 处理非ASCII字符
let japanese = "こんにちは";
println!("日文长度(bytes): {}", japanese.len());
println!("日文字符数: {}", japanese.chars().count());
// 使用 encode_utf16 处理 Unicode
let text = "hello";
let encoded: Vec<u16> = text.encode_utf16().collect();
println!("UTF-16 编码: {:?}", encoded);
// 遍历 Unicode 标量值
for scalar in text.chars() {
println!("字符: {}, Unicode: U+{:04X}", scalar, scalar as u32);
}
// 处理生字符串 (raw string)
let raw = r"C:\Users\name\documents";
println!("原始路径: {}", raw);
// 处理包含引号的字符串
let quoted = "他说 '你好'";
println!("带引号: {}", quoted);
}

今天我们深入学习了Rust的字符串类型:

  • String vs &str:String是拥有的堆分配字符串,&str是借用的字符串引用
  • UTF-8编码:Rust字符串是UTF-8编码,按字节索引而非按字符索引
  • 常用操作:push、push_str、+、format!、切片等
  • 遍历方式:按字符 chars() 或按字节 bytes()
  • 编码注意:处理非ASCII字符时需要使用 chars() 或专门的编码库

字符串处理在日常编程中非常常见,建议多加练习。明天我们将学习文件与输入输出操作。