章节目录

Lowering:把暂停后的计算装进 closure

本节阅读量:

现在来到本章的核心。我们不保存 C++ 调用栈,也不增加 yield 这种 IR 指令;而是把 body 改写成已有语言中的 closure、cell、set!、record 和普通 call。

仍使用这个例子:

1
2
3
(coroutine
  (let x (yield 40)
    (+ x 2)))

它的关键不是“40 被产出了”,而是第一次暂停后,系统仍必须记住:下一次要建立 x = 40,然后计算 (+ x 2),再把结果包装为 done record。

先为每个 coroutine 建立 state

概念上的 core 形状如下。为便于阅读,省略了实际输出中的数字后缀:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
(let $coro.state 0
  (let $coro.finish
    (lambda value
      ... 把 value 变成稳定的 done record ...)
    (begin
      (set! $coro.state initial-step)
      (lambda ignored
        (let next $coro.state
          (begin
            (set! $coro.state 0)
            (next 0)))))))

最外层返回的是 resume closure。它每次做三件事:

  1. 从 state 读出当前 step closure;
  2. 先把 state 写成 0;
  3. 调用刚才读出的 step。

临时写入 0 是本章的重入保护。step 若在自己尚未返回时再次 (resume g),新调用会尝试把整数 0 当作 closure,触发已有的 check-closure 错误,而不会递归重新开始这个 step。

创建 coroutine 时只安装 initial-step,并不调用它,因此 body 保持惰性。

这些绑定名由 pass 统一生成:

1
2
3
std::string fresh(const std::string& role) {
    return "$coro." + role + "." + std::to_string(next_name_++);
}

源码 identifier 只能是 letter+,所以 $coro.state.0 不可能与读者写下的变量重名。所有嵌套 coroutine 共用同一个单调计数器,也不会彼此碰撞。

yield 的变换

对于:

1
(yield value)

lowering 先正常求 value,把结果绑定到内部名字 $coro.yieldvalue。随后创建一个 next step:

1
2
3
4
5
6
(let $coro.yieldvalue value
  (begin
    (set! $coro.state
      (lambda ignored
        continuation($coro.yieldvalue)))
    (record 0 $coro.yieldvalue)))

这里的 continuation(...) 不是新语言语法,而是 lowering 已经构造好的 closure call。它代表 yield 后原本还要执行的程序。

顺序很重要:

1
2
3
先算 value
再创建并安装 next step
最后返回 (record 0 value)

所以 (yield (begin (set! x 1) 40)) 会先改 x,再暂停;而 payload 的错误不会让 state 指向一个并不存在的后续步骤。

对应的 C++ 分支直接构造这棵 core AST:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
case ExprKind::yield: {
    auto& yield = static_cast<YieldExpr&>(*expr);
    const std::string value_name = fresh("yieldvalue");
    const std::string next_argument = fresh("nextarg");

    auto next_step = std::make_unique<LambdaExpr>(
        next_argument,
        call(continuation, variable(value_name)));
    auto pause = std::make_unique<BeginExpr>(
        std::make_unique<SetExpr>(state, std::move(next_step)),
        protocol_record(0, variable(value_name)));
    return std::make_unique<LetExpr>(
        value_name,
        lower_regular(std::move(yield.value)),
        std::move(pause));
}

lower_regular 只处理不能包含当前 yield 的普通表达式;protocol_record(0, ...) 则构造已有的二字段 RecordExpr。这里没有绕过 AST 直接生成 IR 字符串。

let 如何把产出值交回给变量

原始程序:

1
2
(let x (yield 40)
  (+ x 2))

在 yield 后,真正的“下一步”是:

1
2
3
(lambda yielded
  (let x yielded
    (finish (+ x 2))))

第一次 resume 的 step 计算 40、把这个 closure 写进 state、返回 (record 0 40)。第二次 resume 调用刚才写入的 closure,于是 yielded 是已经保存的 40,再建立原来的 x。

这正是“yield 在下一次恢复后作为自己的表达式值继续”的实现。它不是重新执行 (yield 40),也不是从恢复者接收一个新值。

实际输出可用 core 查看:

1
2
cd code/11_coroutines
./mini core examples/coroutine_protocol.lang

其中会看到形如:

1
Lambda($coro.letvalue..., Let(x, Var($coro.letvalue...), ...))

这就是为原始 let x 建立的后续步骤。

begin 和 if 也把“剩余工作”包成 closure

对二元 begin:

1
(begin first second)

lowering 先创建一个“执行 second”的 continuation closure,再让 first 在这个 continuation 下运行。若 first 发生 yield,保存的正是该 closure;若 first 没暂停,它也通过普通 closure call 进入 second。

对 if:

1
(if condition then else)

lowering 创建一个接收 condition value 的 closure。恢复后它按已有的 if 规则选择 then 或 else。两个分支共享同一个外层 continuation,而不是把 yield 后的大段代码分别复制进两个分支。

1
2
3
4
condition 的 continuation
  -> if condition-value
       then 在同一个后续 continuation 下运行
       else 也在同一个后续 continuation 下运行

这条共享很重要。若每遇到一个 if 都把“后面的程序”复制到 then 和 else,嵌套分支会使生成 AST 快速膨胀;绑定一个 continuation closure 后,后续代码只保留一份。

因此 lower_flow 接收的是 continuation 的名字,不是一棵需要复制的 continuation AST:

1
2
3
4
std::unique_ptr<Expr> lower_flow(
    std::unique_ptr<Expr> expr,
    const std::string& state,
    const std::string& continuation);

if 的两个分支都引用同一个 continuation:

1
2
3
4
5
6
7
8
auto lowered_then =
    lower_flow(std::move(then_branch), state, continuation);
auto lowered_else =
    lower_flow(std::move(else_branch), state, continuation);
auto choose_branch = std::make_unique<IfExpr>(
    variable(condition_parameter),
    std::move(lowered_then),
    std::move(lowered_else));

外层再把 choose_branch 放进一个接收 condition value 的 closure,并让 condition 在这个新 continuation 下运行。无论 then 还是 else 被选择,原来的后续代码都只通过同一个变量到达。

finish closure 让 done 结果稳定

body 到达普通结果 value 时,不应直接返回 (record 1 value) 就结束。否则第二次完成后的 resume 会重新分配一个 record,eq? 将失败。

dispatcher 每次调用 step 前都会把 state 暂时写成 0,所以 done step 不能只是返回 completed;它还要在返回前把自己重新装回 state。finish closure 的概念形状是:

1
2
3
4
5
6
7
8
9
(lambda value
  (let completed (record 1 value)
    (letrec done ignored
      (begin
        (set! $coro.state done)
        completed)
      (begin
        (set! $coro.state done)
        completed))))

它只在 body 真正结束时运行一次:先分配 completed record,再建立自引用的 done closure,把 done 写入 state,最后把 completed 交给当前 resume。以后每次 resume 都调用同一个 done;done 先恢复 state,再返回同一个 completed。这样无论恢复多少次,都不会重跑 body、丢失 done step 或重新分配完成 record。

在 C++ 中,自引用不是靠悬空指针拼出来的,而是继续复用已有的 LetRecExpr:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
auto done_body = std::make_unique<BeginExpr>(
    std::make_unique<SetExpr>(state, variable(done)),
    variable(completed));
auto install_done = std::make_unique<BeginExpr>(
    std::make_unique<SetExpr>(state, variable(done)),
    variable(completed));
auto stable_done = std::make_unique<LetRecExpr>(
    done,
    done_argument,
    std::move(done_body),
    std::move(install_done));

第八章以后,letrec 的 self binding 已经会变成共享 cell;第七章 closure 也会捕获 completed 与 state。本章只是在新的组合里使用它们。

coroutine_done_no_restart.lang 检查了另一个同样重要的性质:即使 body 在结束前执行了 set!,反复 resume 也只会让那段副作用发生一次。

为什么 cell、closure 和 record 足够

从第八章的角度,$coro.state 是一个可变绑定,最终会 lower 成 cell;从第七章的角度,每个 continuation 是普通 closure;从第六章的角度,pause/done protocol 是普通二字段 record。

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
state cell
  -> initial step / next step / done step

next step closure
  -> yield value
  -> let/begin/if 的剩余控制流
  -> 外层捕获值

result record
  -> status 0 或 1
  -> yielded value 或 final value

这些对象都已经遵守统一 heap-object header 协议。第九章的 copying collector 不需要知道“这是协程”;它只沿 closure capture、cell payload 和 record field 继续扫描。第十章也不需要新分配策略:可能是堆引用的 state 和 continuation 保守地保存在 root slot,确定为整数的临时计算仍可进入寄存器。

下一节会沿解释器、IR 和原生 GC 路径验证,这个 surface-to-core 变换确实让暂停状态跨越后续分配仍能正确恢复。


11.3 AST 与 parser:先保留表面语法,再统一去糖

上一节

11.5 解释器:执行 core,而不是认识协程

下一节