字符串与 bytes
本章讲解Go字符串处理:字符串不可变性、拼接方法、bytes/rune操作、字符编码。重点:Go字符串是UTF-8字节序列,理解rune(码点)和字节的区别,掌握高效拼接。
前置知识:第3章 语法基础 学习目标:掌握字符串操作、理解UTF-8编码、会处理中文字符
8.1 字符串是不可变的
Section titled “8.1 字符串是不可变的”Go字符串特性
Section titled “Go字符串特性”Go的字符串是只读的字节序列:
s := "hello"// s[0] = 'H' // 编译错误:无法修改字符串
// 必须创建新字符串s2 := "H" + s[1:] // "Hello"Python字符串同样不可变:
s = "hello"s[0] = "H" # TypeErrors = "H" + s[1:] # 正确,创建新字符串字符串内部表示
Section titled “字符串内部表示”// 字符串结构(伪代码)type StringHeader struct { Data unsafe.Pointer // 指向字节数组 Len int // 长度(字节数)}字符串底层是字节数组,UTF-8编码。
字符串与切片互相转换
Section titled “字符串与切片互相转换”str := "hello"
// 字符串转字节切片bytes := []byte(str)
// 字节切片转字符串str2 := string(bytes)
// 注意:每次转换都分配新内存rune类型
Section titled “rune类型”rune是int32的别名,表示Unicode码点:
var r rune = '中' // '中'的码点是U+4E2Dfmt.Printf("%c\n", r) // 中fmt.Printf("%U\n", r) // U+4E2D
// 遍历字符串获取runestr := "Hello世界"for i, r := range str { fmt.Printf("%d: %c (U+%04X)\n", i, r, r)}8.2 字符串拼接方法
Section titled “8.2 字符串拼接方法”本节目的:掌握Go字符串拼接的多种方式及性能差异
s1 := "hello"s2 := "world"s3 := s1 + " " + s2 // "hello world"简单场景适用,复杂场景效率低。
fmt.Sprintf
Section titled “fmt.Sprintf”name := "Alice"age := 30s := fmt.Sprintf("Name: %s, Age: %d", name, age)// "Name: Alice, Age: 30"用于格式化复杂场景。
strings.Join
Section titled “strings.Join”import "strings"
parts := []string{"hello", "world", "go"}s := strings.Join(parts, " ") // "hello world go"
// 常用场景:路径拼接paths := []string{"dir1", "dir2", "file.txt"}s := strings.Join(paths, "/") // "dir1/dir2/file.txt"strings.Builder
Section titled “strings.Builder”import "strings"
var builder strings.Builder
builder.WriteString("hello")builder.WriteString(" ")builder.WriteString("world")
result := builder.String() // "hello world"
// 写入格式化内容fmt.Fprintf(&builder, "Name: %s, Age: %d", name, age)拼接性能对比
Section titled “拼接性能对比”// 效率低:不推荐循环内使用s := ""for _, item := range items { s += item // 每次都创建新字符串}
// 效率高:推荐var parts []stringfor _, item := range items { parts = append(parts, item)}s := strings.Join(parts, "")
// 或使用Buildervar builder strings.Builderfor _, item := range items { builder.WriteString(item)}s := builder.String()Python对比
Section titled “Python对比”# 加号s = "hello" + " " + "world"
# join方法parts = ["hello", "world", "go"]s = " ".join(parts) # "hello world go"
# f-strings = f"Name: {name}, Age: {age}"8.3 bytes包与rune
Section titled “8.3 bytes包与rune”本节目的:掌握bytes包操作字节切片、理解rune(Unicode码点)的处理
bytes包
Section titled “bytes包”处理字节切片的工具库:
import "bytes"
// Buffer:动态增长的字节切片var buf bytes.Bufferbuf.Write([]byte("hello"))buf.WriteByte(',')buf.WriteString(" world")
data := buf.Bytes() // []bytestr := buf.String() // string
// Reader:读取字节切片reader := bytes.NewReader([]byte("hello world"))b, _ := reader.ReadByte() // 'h'
// 常用函数bytes.Split([]byte("a,b,c"), []byte(",")) // [][]byte{a, b, c}bytes.Contains([]byte("hello"), []byte("ll")) // truebytes.Count([]byte("hello"), []byte("l")) // 2bytes.TrimSpace([]byte(" hello ")) // []byte("hello")strings包常用函数
Section titled “strings包常用函数”import "strings"
// 大小写strings.ToLower("HELLO") // "hello"strings.ToUpper("hello") // "HELLO"strings.Title("hello world") // "Hello World"
// 查找strings.Index("hello", "ll") // 2strings.Contains("hello", "ll") // truestrings.HasPrefix("hello", "he") // truestrings.HasSuffix("hello", "lo") // true
// 分割fields := strings.Fields(" hello world ") // ["hello", "world"]parts := strings.Split("a,b,c", ",") // ["a", "b", "c"]parts = strings.SplitN("a,b,c", ",", 2) // ["a", "b,c"]
// 替换strings.Replace("hello world", "world", "go", 1) // "hello go"strings.ReplaceAll("aaa", "a", "b") // "bbb"
// 修整strings.Trim(" hello ", " ") // "hello"strings.TrimLeft(" hello", " ") // "hello"strings.TrimRight("hello ", " ") // "hello"strings.TrimPrefix("hello", "he") // "llo"strings.TrimSuffix("hello", "lo") // "hel"rune与字符串操作
Section titled “rune与字符串操作”str := "Hello世界"
// 统计字符数(rune数)fmt.Println(len([]rune(str))) // 8
// 提取子串(按字符)func substringByRune(s string, start, end int) string { runes := []rune(s) return string(runes[start:end])}
sub := substringByRune(str, 6, 8) // "世界"
// 反转字符串func reverseString(s string) string { runes := []rune(s) for i, j := 0, len(runes)-1; i < j; i, j = i+1, j-1 { runes[i], runes[j] = runes[j], runes[i] } return string(runes)}8.4 字符编码处理
Section titled “8.4 字符编码处理”本节目的:理解UTF-8编码规则、掌握rune遍历和中文字符处理
UTF-8编码
Section titled “UTF-8编码”Go字符串默认是UTF-8编码:
str := "Hello世界"fmt.Println(len(str)) // 12(字节),5个ASCII + 3字节(世) + 3字节(界)
// UTF-8编码规则:// ASCII: 1字节 (0-127)// 其他: 2-4字节,以特定模式区分
// 验证UTF-8valid := utf8.ValidString(str) // trueimport ( "golang.org/x/text/encoding/simplifiedchinese" "golang.org/x/text/transform")
// GBK转UTF-8gbkData := []byte{...}reader := transform.NewReader(bytes.NewReader(gbkData), simplifiedchinese.GBK.NewDecoder())utf8Data, _ := io.ReadAll(reader)import "unicode"
c := 'A'
unicode.IsLetter(c) // trueunicode.IsDigit(c) // falseunicode.IsUpper(c) // trueunicode.IsLower(c) // falseunicode.IsSpace(c) // falseunicode.IsPunct(c) // false
// 转换unicode.ToUpper(c) // 'A'unicode.ToLower(c) // 'a'
// 判断中文字符func isChinese(r rune) bool { return unicode.Han(r)}
for _, r := range "Hello世界" { fmt.Printf("%c: %v\n", r, isChinese(r))}// H: false// e: false// l: false// l: false// o: false// 世: true// 界: true字符串与数字互转
Section titled “字符串与数字互转”import "strconv"
// int转strings := strconv.Itoa(123) // "123"s := strconv.FormatInt(123, 10) // "123"s := strconv.FormatInt(123, 2) // "1111011"
// string转inti, _ := strconv.Atoi("123") // 123i, _ := strconv.ParseInt("123", 10, 64) // 123
// 浮点数转strings := strconv.FormatFloat(3.14, 'f', 2, 64) // "3.14"s := fmt.Sprintf("%.2f", 3.14) // "3.14"
// string转浮点数f, _ := strconv.ParseFloat("3.14", 64) // 3.14JSON与字符串
Section titled “JSON与字符串”import "encoding/json"
// 字符串需要转义data := map[string]string{"key": "value\nwith\nnewlines"}jsonBytes, _ := json.Marshal(data)// {"key":"value\nwith\nnewlines"}
// 解析时自动还原- 频繁拼接用strings.Builder:循环内拼接字符串用
strings.Builder,避免+号每次创建新字符串 - 遍历中文字符用rune:用
for i, r := range str遍历,正确处理多字节UTF-8字符 - 字节vs字符:len(str)是字节数,不是字符数;中文字符在UTF-8占3字节
- JSON字符串转义:Go的encoding/json会自动处理换行符、双引号等特殊字符
- 转换开销:string和[]byte互相转换有内存分配开销,避免在热路径中频繁转换
| Python | Go |
|---|---|
str不可变 | string不可变 |
s.encode('utf-8') | []byte(s) |
b.decode('utf-8') | string(b) |
" ".join(list) | strings.Join(list, " ") |
f"Name: {n}" | fmt.Sprintf("Name: %s", n) |
str(x) | strconv.Itoa(x) |
int(s) | strconv.Atoi(s) |
| Unicode原生 | rune表示Unicode码点 |
- 实现字符串反转函数
- 统计字符串中中文字符的数量
- 实现一个函数,将”hello_world”转换为”helloWorld”
- 用strings.Builder拼接1000个字符串
- 实现字符串截取函数,支持UTF-8字符边界