章节目录

汇编:在通用堆对象里放入代码和环境

本节阅读量:

IR 已经说明“创建 closure”和“检查并调用 closure”,却没有规定对象占多少字节。汇编 emitter 要把这些逻辑操作落实为 header、payload、tagged pointer、调用寄存器和间接 call。

本节只讲 closure 带来的新机器表示。整数编码和 record 的一般布局沿用第六章,不再从头展开。

共用一个 heap-object pointer tag

运行时常量集中在 runtime/value.h:

1
2
3
4
5
6
7
8
inline constexpr long kTagMask = 7;
inline constexpr long kIntTag = 1;
inline constexpr long kHeapObjectTag = 4;

enum class HeapKind : long {
    record = 1,
    closure = 2,
};

record 和 closure 引用的机器字低三位都为 4:

1
tagged heap Value = aligned raw object address | 4

具体种类由 raw object 的 header 低八位决定。检查低位 tag 只能说明“这是某种堆对象”,不能说明它一定是 closure。

closure 复用 header 公式

第六章已经定义:

1
2
header = (payload_count << 8) | heap_kind
object_size = 8 * (1 + payload_count)

closure 的 payload 由一个 raw code pointer 和若干捕获 Value 组成,因此:

1
2
payload_count = 1 + capture_count
closure_size  = 8 * (2 + capture_count)

辅助函数直接复用通用协议:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
inline constexpr std::size_t closure_header(
        std::size_t capture_count) {
    return heap_header(HeapKind::closure, capture_count + 1);
}

inline constexpr std::size_t closure_size(
        std::size_t capture_count) {
    return heap_object_size(capture_count + 1);
}

inline constexpr std::size_t closure_code_offset() {
    return heap_payload_offset(0);
}

inline constexpr std::size_t closure_capture_offset(
        std::size_t index) {
    return heap_payload_offset(index + 1);
}

closure 的准确布局

raw object 从 header 开始:

1
2
3
4
5
offset 0       header: kind = closure, count = 1 + n
offset 8       raw code pointer
offset 16      capture 0: tagged Value
offset 24      capture 1: tagged Value
...

code pointer 是汇编标签地址,不是源语言 Value;捕获槽保存的则是完整 tagged Value。header kind 让后续运行时能够按对象种类解释这些不同 payload。

三个常用大小是:

捕获数 payload count header 对象大小
0 1 `(1 « 8) 2 = 258`
1 2 `(2 « 8) 2 = 514`
2 3 `(3 « 8) 2 = 770`

即使没有捕获值,closure 也必须保存 header 和 code pointer,所以最小对象仍是 16 字节。

为什么不能继续使用裸代码地址

第六章之前,无捕获函数可以只把代码地址编码成一个函数值。第七章把所有函数统一成 closure:

1
2
无捕获 lambda   closure(label, [])
有捕获 lambda   closure(label, [value ...])

调用端因此只保留一条协议,不需要先猜一个函数值究竟是裸地址还是对象指针。原来的函数专用 pointer tag 也不再需要。

创建一个捕获 x 的 closure

closure.lang 捕获一个值,所以 emitter 申请 24 字节。在 macOS 目标上,关键汇编是:

1
2
3
4
5
6
7
8
9
movq $24, %rdi
call _malloc
movq $514, 0(%rax)
leaq .Llambda0(%rip), %rcx
movq %rcx, 8(%rax)
movq -8(%rbp), %rcx
movq %rcx, 16(%rax)
orq $4, %rax
movq %rax, -16(%rbp)

执行顺序是:

  1. malloc 返回 raw object address;
  2. offset 0 写 closure header;
  3. offset 8 写 raw code pointer;
  4. offset 16 写捕获的 tagged x;
  5. 最后把通用 heap tag 4 加到返回 Value 上。

Linux 目标使用 malloc 标签,macOS 使用 _malloc;对象布局完全相同。

捕获值已经由 IR 固定

OpKind::closure 的 fields 按稳定顺序保存 operands。emitter 调用 malloc 后,再从这些 operand 的栈槽逐个装载 Value:

1
2
3
4
5
6
7
for (std::size_t i = 0; i < op.fields.size(); ++i) {
    out_ += "    movq " + operand_text(op.fields[i]) +
            ", %rcx\n";
    out_ += "    movq %rcx, " +
            std::to_string(closure_capture_offset(i)) +
            "(%rax)\n";
}

汇编层不再做自由变量分析,也不根据源码名字排序。它严格消费 IR 给出的顺序。

为什么检查要成为独立 IR op

源语言规定 call 的顺序是:先求 callee,确认它是 closure,然后才求 argument。若只在最终 call op 中检查,argument 的 lowering ops 已经排在前面,错误程序就可能先执行 argument。

因此结构化 IR 明确插入:

1
2
3
4
callee = ...
check-closure callee
argument = ...
result = call callee argument

OpKind::closure_check 不产生新 Value,却保留一项可观察的运行时检查顺序。汇编 emitter 在这里生成 tag 和 header kind guard。

closure check 先看 tag,再读 header

检查逻辑是:

1
2
3
4
5
6
7
8
callee low tag != 4
    -> runtime error

callee low tag == 4
    -> 清 tag 得到 raw object address
    -> 读取 header kind
    -> kind != closure
       -> runtime error

对应代码形状是:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
movq callee_slot, %rax
movq %rax, %rcx
andq $7, %rcx
cmpq $4, %rcx
jne .Lruntime_error
andq $-8, %rax
movq 0(%rax), %rdx
andq $255, %rdx
cmpq $2, %rdx
jne .Lruntime_error

顺序不能交换。整数的 tagged word 不是有效地址,必须通过 low tag 检查以后才能解引用 header。

record 同样拥有 low tag 4,但 kind 为 1,所以会在第二层检查失败。生成代码不会把 record payload 当成代码入口。

调用约定携带隐藏 closure 参数

第七章的机器调用约定是:

1
2
3
%rdi  tagged closure Value
%rsi  tagged user argument Value
%rax  tagged return Value

%rdi 传 tagged Value,而不是 raw pointer。这一点让递归 self 和捕获到的对象引用始终保持统一 Value 形状。

closure_check 已经验证 callee。最终 call op 只需重新取出 code pointer:

1
2
3
4
5
movq callee_slot, %rdi
movq %rdi, %rcx
andq $-8, %rcx
movq argument_slot, %rsi
call *8(%rcx)

这里 %rcx 是临时 raw address,%rdi 仍保留原始 tagged closure Value。间接 call 从 raw object 的 offset 8 读取代码入口。

函数入口怎样读取参数和捕获值

函数体进入时,emitter 为参数和捕获 locals 分配栈槽。普通闭包先保存 %rsi 参数,再用 %rdi 的副本读取 capture:

1
2
3
4
5
movq %rsi, -8(%rbp)
movq %rdi, %rax
andq $-8, %rax
movq 16(%rax), %rcx
movq %rcx, -16(%rbp)

此后函数体里的名字都统一从 IR local 栈槽读取:

1
2
参数 y2      -8(%rbp)
捕获 x1      -16(%rbp)

closure_capture_offset(0) 是 16,不是 8;offset 8 已经被 code pointer 占用。

self 来自 %rdi,不占 payload

递归函数的 IrFunction::self 非空时,入口额外保存:

1
movq %rdi, self_slot

以 recursive_closure.lang 为例,closure 只捕获 base,所以仍然是:

1
2
3
header       514
object size  24
capture 0    base

它不会因为递归而变成两个 capture。递归分支执行:

1
call adddown1 next_argument

其中 adddown1 就是入口从 %rdi 保存的 tagged actual callee。call 又把它放回 %rdi,形成递归。

self 与参数同名不会共用栈槽

IR 已经把 self 和 param alpha-rename,例如:

1
2
lambda0(f1, self: f0, captures: []):
  return f1

栈槽分配器按这两个不同名字分别分配位置。入口先保存 %rdi 到 f0,再保存 %rsi 到 f1。返回使用 f1,因此参数正确遮蔽 self。

closure 和 record 可以互相引用

closure 捕获 record 时,capture 槽保存 tagged record pointer:

1
2
3
4
5
6
closure object
  header kind 2
  code pointer
  capture 0 --------> record object
                       header kind 1
                       fields...

record 保存 closure 时,record field 保存 tagged closure pointer:

1
2
3
4
5
record object
  header kind 1
  field 0 -----------> closure object
                       header kind 2
                       code pointer

两者都只复制一个机器字引用,不内联另一个对象。访问时再由具体操作检查 header kind。

record? 为什么能安全拒绝 closure

record? 看到 low tag 4 后还会读取 header kind:

1
2
kind == record   -> tagged 1
kind != record   -> tagged 0

所以:

1
(record? (lambda x x))

返回 0,不会报错。相反,size 和 get 要求 operand 必须是 record,收到 kind 2 时会进入共享错误出口。

eq? 显式拒绝 closure

机器字地址相同并不等于源语言定义了函数相等。emit_equal 在普通比较前分别检查两个 operand:若它是 heap object 且 header kind 为 closure,就跳到错误出口。

整数和 record 继续使用原有比较规则;closure 则不参与 eq?。这与解释器的错误语义一致。

共享运行时错误出口

本章这些编译后错误都会跳到同一个标签:

1
2
3
4
call 的 callee 不是 closure
eq? 的任一 operand 是 closure
size/get 的 operand 不是 record
get 的 index 越界

错误出口是:

1
2
3
4
.Lruntime_error:
    movl $70, %edi
    call _exit
    ud2

Linux 使用 exit,macOS 使用 _exit。只有程序 IR 中实际出现需要错误路径的操作时,emitter 才追加这段代码。

解释器提供具体错误文字;生成程序只用退出状态 70 区分运行时失败。本章没有引入异常对象或诊断字符串 runtime。

固定栈帧保持调用对齐

汇编 emitter 先为函数的 self、参数、捕获 locals 和所有 op 结果分配栈槽,再把帧大小向上对齐到 16 字节。prologue 形状是:

1
2
3
pushq %rbp
movq %rsp, %rbp
subq $frame_size, %rsp

frame_size 是 16 的倍数,因此调用 malloc、exit 或 closure code 时满足 x86-64 System V ABI 的栈对齐要求。

手动查看生成结果

生成汇编:

1
2
cd code/07_closures
./mini compile examples/closure.lang -o out.s

Linux、WSL 和 Intel Mac 手动链接:

1
2
3
cc out.s -o out
./out
echo $?

Apple Silicon Mac 使用:

1
2
3
cc -arch x86_64 out.s -o out
./out
echo $?

正常结果为 42。若改为编译 call_record_error.lang、function_eq_error.lang、record_size_closure_error.lang 或 record_get_closure_error.lang,退出状态应为 70。

当前内存边界

closure 和 record 仍通过 malloc 分配,不调用 free。本章的目标是建立正确的对象种类、捕获布局和调用约定,不在这里加入回收器。

统一 header 已经说明对象有多少 payload word,closure kind 也说明第一个 payload 是 raw code pointer;这些信息会成为后续对象扫描的基础,但本章不提前实现 GC。


7.4 结构化 IR:把代码标签和捕获值配成一对

上一节

7.6 闭包:从一行源码走到一次间接调用

下一节