Skip to content

第8章 标准类型转换

pybind11 的核心能力之一是自动类型转换。当 Python 代码调用 C++ 函数时,pybind11 会自动将 Python 对象转换为对应的 C++ 类型;返回值则反向转换。本章详细剖析这一过程。

Python 的 True/False 与 C++ 的 bool 类型无缝对应:

#include <pybind11/pybind11.h>
namespace py = pybind11;
bool negate(bool value) {
return !value;
}
PYBIND11_MODULE(example, m) {
m.def("negate", &negate, "Negates a boolean value");
}
>>> import example
>>> example.negate(True)
False
>>> example.negate(False)
True

原理:pybind11 使用 type_caster<bool> 完成双向转换。非零值视为 True,零视为 False。

pybind11 支持所有标准整数类型:

m.def("add_int", [](int8_t a, int8_t b) { return a + b; });
m.def("add_int", [](int16_t a, int16_t b) { return a + b; });
m.def("add_int", [](int32_t a, int32_t b) { return a + b; });
m.def("add_int", [](int64_t a, int64_t b) { return a + b; });
m.def("add_uint", [](uint8_t a, uint8_t b) { return a + b; });
m.def("add_uint", [](uint16_t a, uint16_t b) { return a + b; });
m.def("add_uint", [](uint32_t a, uint32_t b) { return a + b; });
m.def("add_uint", [](uint64_t a, uint64_t b) { return a + b; });

Python 的 int 是任意精度,转换时若值超出目标类型范围会抛出 OverflowError。

float 和 double 直接对应 Python 的 float(实际上是 C++ double 的双精度):

double average(double a, double b) {
return (a + b) / 2.0;
}
>>> example.average(3.5, 7.2)
5.35

关键洞察:Python 的 float 类型等价于 IEEE 754 双精度浮点数,精度约 15-17 位十进制数字。C++ 的 float 单精度在 Python 端会被提升为 double。

std::string 与 Python str 之间的转换是自动的,但有一个重要前提:UTF-8 编码。

#include <pybind11/pybind11.h>
#include <string>
std::string greet(const std::string& name) {
return "Hello, " + name + "!";
}
PYBIND11_MODULE(example, m) {
m.def("greet", &greet);
}
>>> example.greet("World")
'Hello, World!'
>>> example.greet("中文")
'Hello, 中文!'

pybind11 默认假设 std::string 持有 UTF-8 数据。如果你的 C++ 代码使用其他编码(如 GBK、Latin-1),需要显式处理:

// 错误示例:假设 source 是 GBK 编码
std::string source = "\xc4\xe3\xba\xc3"; // "你好" 的 GBK 编码
// 正确做法:显式转换或使用 py::bytes
py::bytes gbk_bytes(source); // 不做编码转换

关键洞察:pybind11 的 type_caster<std::string> 假设 UTF-8 编码。这对现代 Linux/macOS 环境是标准做法,但 Windows 传统上使用 UTF-16。如果你的字符串可能包含非 ASCII 字符,确保 C++ 端也使用 UTF-8。

Python 3 的 str 是 Unicode 对象。在 pybind11 中:

// C++ 接收 Unicode 字符串
void print_unicode(const std::string& s) {
std::cout << "Received: " << s << std::endl;
}
// 返回 Unicode 字符串
std::string get_unicode() {
return u8"你好世界"; // UTF-8 编码的中文
}
>>> example.print_unicode("你好")
Received: 你好
>>> example.get_unicode()
'你好世界'

如果需要更精细的 Unicode 控制,可使用 py::str:

py::str py_str = py::cast("你好"); // 显式创建 Python str 对象
m.attr("unicode_value") = py_str;

std::string 不仅可以表示文本,也可以表示字节序列。pybind11 区分两者:

C++ 类型Python 类型说明
std::stringstr文本,假设 UTF-8
py::bytesbytes原始字节序列
#include <pybind11/pybind11.h>
namespace py = pybind11;
// 处理字节序列
py::bytes process_bytes(py::bytes input) {
std::string data = input.cast<std::string>();
// 处理字节...
return py::bytes(data);
}
PYBIND11_MODULE(example, m) {
m.def("process_bytes", &process_bytes);
}
>>> data = b'\x00\x01\x02\x03'
>>> result = example.process_bytes(data)
>>> result
b'\x00\x01\x02\x03'

关键洞察:区分「文本」和「字节」是 Python 3 的核心理念。pybind11 通过 std::string vs py::bytes 延续这一区分。如果你不确定数据编码,始终使用 bytes 并在 C++ 端显式解码。

Python 的 None 在 C++ 中对应 py::none 类型:

#include <pybind11/pybind11.h>
namespace py = pybind11;
bool is_empty(const py::object& obj) {
return obj.is_none();
}
py::object return_none() {
return py::none();
}
PYBIND11_MODULE(example, m) {
m.def("is_empty", &is_empty);
m.def("return_none", &return_none);
}
>>> example.is_empty(None)
True
>>> example.is_empty("hello")
False
>>> example.return_none() is None
True

关键洞察:py::none 是一个特殊类型,表示 Python 的 None 值。在函数签名中使用 py::object 可以接收任意类型,然后检查 is_none()。如果直接使用 py::none& 作为参数类型,则该参数必须为 None,否则抛出 TypeError。

Python 类型C++ 类型转换方向
boolbool双向自动
intint8_t ~ int64_t双向自动(超范围抛异常)
intuint8_t ~ uint64_t双向自动(超范围抛异常)
floatfloat, double双向自动
strstd::string双向(假设 UTF-8)
bytespy::bytes双向(原始字节)
Nonepy::none双向

关键洞察:pybind11 的自动类型转换基于 type_caster 模板机制。每种支持类型的转换器都定义了 load() 和 cast() 方法。这一机制是可扩展的——你可以通过特化 type_caster 来支持自定义类型(详见第 15 章)。