章节目录

汇编:分配可变大小对象,并在读取前检查

本节阅读量:

汇编 emitter 只消费结构化 IrProgram。对于:

1
(get (record 40 2) 1)

它接收的是:

1
2
3
t.0 = record 40 2
t.1 = get t.0 1
return t.1

这一层才把逻辑操作落实为 tagged word、malloc、header、payload 写入、种类检查、边界检查和内存读取。

运行时常量集中在一个地方

src/runtime/value.h 定义所有表示公式:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
inline constexpr long kTagMask = 7;
inline constexpr long kIntTag = 1;
inline constexpr long kFunctionTag = 3;
inline constexpr long kHeapObjectTag = 4;

inline constexpr long kWordSize = 8;
inline constexpr long kHeaderKindMask = 0xff;
inline constexpr int kHeaderPayloadCountShift = 8;
inline constexpr int kRuntimeErrorExitCode = 70;

enum class HeapKind : long {
    record = 1,
};

辅助函数把通用公式写成可复用接口:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
inline constexpr std::size_t heap_header(HeapKind kind,
                                         std::size_t payload_count) {
    return (payload_count << kHeaderPayloadCountShift) |
           static_cast<std::size_t>(kind);
}

inline constexpr std::size_t heap_object_size(
        std::size_t payload_count) {
    return (payload_count + 1) * kWordSize;
}

inline constexpr std::size_t heap_payload_offset(
        std::size_t index) {
    return (index + 1) * kWordSize;
}

record 的三个辅助函数只是给通用协议加上 HeapKind::record:

1
2
3
record_header(n)       = heap_header(record, n)
record_size(n)         = heap_object_size(n)
record_field_offset(i) = heap_payload_offset(i)

后续 closure 可以复用 heap_header、heap_object_size 和 heap_payload_offset,但拥有自己的 header kind 与 payload 解释。本章不提前实现 closure。

tag 与 header 必须分开理解

一个 record 在运行时同时有两层标记:

1
2
3
4
5
6
7
栈槽中的 Value
    raw_address | 4
    4 表示 heap object

raw object 的 offset 0
    (field_count << 8) | 1
    1 表示 HeapKind::record

低位 4 不是 record 专用 tag。只检查 4,未来会把 closure 或 cell 也误认为 record;因此 record?、size 和 get 还必须检查 header kind。

旧值操作适配 tagged integer

统一表示后,整数 operand 在汇编中写成:

1
2
3
case OperandKind::integer:
    return "$" +
           std::to_string(tagged_integer(operand.integer_value));

所以 IR 的 40 变成汇编立即数 321。

加法对两个 8n + 1 编码相加后减去多余的一个 tag:

1
2
3
movq lhs, %rax
addq rhs, %rax
subq $1, %rax

sub1 减少一个源语言整数,机器编码应减 8:

1
subq $8, %rax

eq? 比较完整机器字。相同整数编码相等,同一个 heap object 的 tagged pointer 也相等;分别分配的对象地址不同。CPU 产生的裸 0/1 还要编码回源语言整数:

1
2
3
4
5
cmpq rhs, %rax
sete %al
movzbq %al, %rax
shlq $3, %rax
orq $1, %rax

函数值仍是 code_address | 3,调用前用 andq $-8 恢复代码地址。这部分沿用第五章。

栈槽保护跨 malloc 的字段 Value

malloc 遵循 System V ABI,可以改写 caller-saved 寄存器。record op 的字段 operand 在 IR 中已经准备好:复杂字段结果位于临时栈槽,简单字段是立即数。

emitter 因此先调用 malloc,再从 operand 的稳定位置逐个装入 %rcx 并写入新对象。它不会要求字段值跨调用一直留在 %rax 或 %rcx。

局部栈大小继续向 16 字节对齐,确保发出 call malloc、call exit 或函数调用时满足 ABI 对齐要求。

根据字段数发出分配

record emitter 不再使用固定 24 和 513。它直接读取 op.fields.size():

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
void emit_record(const Op& op) {
    out_ += "    movq $" +
            std::to_string(record_size(op.fields.size())) +
            ", %rdi\n";
    out_ += "    call " + malloc_label(target_) + "\n";
    out_ += "    movq $" +
            std::to_string(record_header(op.fields.size())) +
            ", 0(%rax)\n";

    for (std::size_t i = 0; i < op.fields.size(); ++i) {
        out_ += "    movq " + operand_text(op.fields[i]) +
                ", %rcx\n";
        out_ += "    movq %rcx, " +
                std::to_string(record_field_offset(i)) +
                "(%rax)\n";
    }

    out_ += "    orq $" +
            std::to_string(kHeapObjectTag) + ", %rax\n";
    store_result(op.dst);
}

这里的循环发生在编译器运行时:字段数量已写在 IR 中,所以 emitter 为每个字段生成一条固定偏移 store。目标程序运行时不需要循环创建这个特定 record。

空 record 仍然调用 malloc

IR:

1
t.0 = record

生成的核心汇编是:

1
2
3
4
movq $8, %rdi
call malloc
movq $1, 0(%rax)
orq $4, %rax

字段循环执行零次,但 header、分配和 tag 都保留。两次 (record) 会执行两次 malloc(8),得到两个不同对象身份。

不能用常量 4 直接表示空 record:那不是有效的 tagged heap address,record? 随后的 header 读取也会访问错误位置。

三字段 record 的生成过程

对于:

1
(record 10 20 30)

字段数为三,因此:

1
2
size   = 8 * (1 + 3) = 32
header = (3 << 8) | 1 = 769

核心汇编形状是:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
movq $32, %rdi
call malloc
movq $769, 0(%rax)

movq $81, %rcx
movq %rcx, 8(%rax)
movq $161, %rcx
movq %rcx, 16(%rax)
movq $241, %rcx
movq %rcx, 24(%rax)

orq $4, %rax

81、161、241 分别是整数 10、20、30 的 tagged 表示。header 中的 count 是裸 3,payload 中的字段则必须是完整 tagged Value。

record? 先看 tag,再看 header kind

predicate 对任意 Value 都应该安全返回,因此不能直接清 tag 并解引用。生成逻辑是:

1
2
3
4
5
6
7
value low tag != heap-object tag
    -> false

value low tag == heap-object tag
    -> 清 tag
    -> 读取 header kind
    -> kind == record 才为 true

对应关键汇编:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
movq value, %rax
movq %rax, %rcx
andq $7, %rcx
cmpq $4, %rcx
jne .Lrecord_predicate_false0

andq $-8, %rax
movq 0(%rax), %rcx
andq $255, %rcx
cmpq $1, %rcx
jne .Lrecord_predicate_false0

movq $9, %rax
jmp .Lrecord_predicate_done0
.Lrecord_predicate_false0:
movq $1, %rax
.Lrecord_predicate_done0:

源语言真值 1 编码为 9,假值 0 编码为 1。内部标签带递增编号,避免一个程序中出现多个 record? 时重名。

即使当前只有 record 一种 heap kind,第二层检查也不是多余工作;它把接口固定在未来能够安全扩展的形式上。

size 与 get 共用 record guard

二者都要求输入一定是 record,因此 emitter 提取共用检查:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
void emit_record_guard(const Operand& operand) {
    runtime_error_used_ = true;
    out_ += "    movq " + operand_text(operand) + ", %rax\n";
    out_ += "    movq %rax, %rcx\n";
    out_ += "    andq $7, %rcx\n";
    out_ += "    cmpq $4, %rcx\n";
    out_ += "    jne .Lruntime_error\n";
    out_ += "    andq $-8, %rax\n";
    out_ += "    movq 0(%rax), %rcx\n";
    out_ += "    movq %rcx, %rdx\n";
    out_ += "    andq $255, %rdx\n";
    out_ += "    cmpq $1, %rdx\n";
    out_ += "    jne .Lruntime_error\n";
}

检查成功后形成一个小约定:

1
2
%rax = untagged raw object address
%rcx = full object header

后续两个 emitter 可以直接继续使用。

最关键的安全顺序是先检查低位 tag,再读取 0(%rax)。整数 42 的机器编码不是有效堆地址,绝不能先解引用再判断 header。

size 解出 count,再编码成 Value

header 的高位保存裸 payload count:

1
shrq $8, %rcx

但源语言 size 要返回 tagged integer,因此完整过程是:

1
2
3
4
shrq $8, %rcx
movq %rcx, %rax
shlq $3, %rax
orq $1, %rax

对于空 record,header 是 1:右移八位得到裸 0,再编码得到机器字 1。主函数最终解码后返回退出状态 0。

get 检查 index 小于 count

guard 成功后,%rcx 保存 header。field emitter 先取得 count:

1
2
3
4
shrq $8, %rcx
movq $index, %rdx
cmpq %rdx, %rcx
jbe .Lruntime_error

AT&T 语法的 cmpq %rdx, %rcx 计算 %rcx - %rdx。这里 %rcx 是 count,jbe 在 count <= index 时跳转,所以只有 index < count 才能继续。

这条比较使用无符号条件,因为 index 和 count 都是非负数量。

检查通过后,用基址、比例索引和 header 偏移一次完成地址计算:

1
movq 8(%rax,%rdx,8), %rax

地址公式是:

1
raw_address + 8 + index * 8

这与 heap_payload_offset(index) = (index + 1) * 8 完全一致。读取结果已经是字段中保存的 tagged Value,不应再次添加 tag。

统一失败出口返回 70

size 类型错误、get 类型错误和 get 越界都跳到:

1
2
3
4
.Lruntime_error:
movl $70, %edi
call exit
ud2

这里调用 C 库 exit(70),而不是简单把 %rax 设为 70 后从当前位置 ret。错误可能发生在任意生成函数中,直接 ret 只会返回调用者并继续执行;exit 能从任何深度终止整个进程。

ud2 表明 exit 理论上不应返回,也防止控制流意外落入后续内容。

解释器仍提供具体错误文字;生成程序只用状态 70 表示本章这几种运行时失败。课程没有在这里引入异常对象或完整诊断 runtime。

guard 同时把 runtime_error_used_ 设为 true。emit() 只有看到这个标志时才追加上述共享错误出口;因此这个判断也覆盖生成的 lambda 函数体,而只含整数、record? 或 record 创建的程序不会携带一段永远用不到的 exit 调用。

平台符号差异

生成代码会按程序实际用到的操作引用 C 库函数:

1
2
malloc    分配对象
exit      终止 size/get 检查失败的程序

只创建 record 的程序只需要 malloc;出现 size 或 get 时,emitter 才生成共享的错误出口并引用 exit。这样没有运行时检查的程序不会凭空多出一个外部符号依赖。

Linux 目标使用 malloc、exit,macOS 目标使用 _malloc、_exit。main 同样分别生成 main 与 _main。

CLI 只生成汇编:

1
./mini compile examples/record_get.lang -o out.s

手动链接时,cc 才从 C 标准库解析这些外部符号:

1
cc out.s -o out

Apple Silicon Mac 使用:

1
cc -arch x86_64 out.s -o out

用退出状态验证成功与失败

成功读取:

1
2
3
4
./mini compile examples/record_get.lang -o out.s
cc out.s -o out
./out
echo $?

record_get.lang 得到 2,退出状态应为 2。

对越界或非 record 的 size/get 示例做相同操作时,退出状态应为 70。shell 只保留较小的进程退出状态,因此本书仍使用小整数作为最终可观察结果。

若顶层直接返回 record,main 检查结果 tag 后返回 0,不会打印对象。完整的 (record ...) 文本只由解释器的 format_value 提供。

嵌套 record 是多个独立分配

对于:

1
(record 0 (record 40 2))

IR 先发出内层 record,再发出外层 record,所以生成代码执行两次 malloc。外层 field 1 保存的是:

1
inner_raw_address | 4

不是把内层 header 与字段复制到外层 payload。对象图在机器层仍由一个机器字引用连接。

本章汇编边界

第六章后端现在实现:

1
2
3
4
5
6
7
8
variable-size allocation
computed header
ordered payload stores
generic heap-object tag
record header-kind check
payload-count extraction
bounds check before field load
shared exit(70) failure path

它仍然不检查 malloc 是否返回空指针,不回收对象,也没有 GC。对旧的算术、函数调用和函数相等,也没有在本章补齐所有动态类型错误路径。

这些边界不妨碍本章目标:解释器和生成代码已经对 record 的创建、判断、长度、安全读取、嵌套与身份实现同一套行为,并把通用 heap object 协议准备好供后续章节扩展。


6.5 结构化 IR:保留一组有序字段

上一节

6.7 边界:这套 record 提供什么,不提供什么

下一节