章节目录

AST 与 parser:先保留表面语法,再统一去糖

本节阅读量:

前几章里,parser 的任务一直是把源码的括号结构转换为 AST。本章新增的三种形式也先保留在 AST 中,而不是在 parser 时立刻变成一大段 closure。

这样做有两个直接好处:

  1. mini ast 仍能显示读者写下的 coroutine、yield、resume;
  2. “yield 是否位于允许的控制流中”可以由一个专门的 validation pass 统一判断。

三个新的 AST 节点

AST 的表达式 kind 增加:

1
2
3
4
5
6
enum class ExprKind {
    // 前十章已有的 kind ...
    coroutine,
    yield,
    resume,
};

对应节点都只有一个子表达式:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
struct CoroutineExpr final : Expr {
    std::unique_ptr<Expr> body;
};

struct YieldExpr final : Expr {
    std::unique_ptr<Expr> value;
};

struct ResumeExpr final : Expr {
    std::unique_ptr<Expr> coroutine;
};

它们和 Sub1Expr、RecordPredicateExpr 一样,使用显式 kind + switch。这意味着 AST 的消费者不能只看 C++ 类型猜测语义;每个 pass 都必须决定是否认识这三个 surface kind。

parser 把它们当作固定特殊形式

普通用户 identifier 依旧只能是 letter+。coroutine、yield、resume 与 eq?、set!、sub1 一样,只在列表 head 的固定分支中识别:

1
2
3
(coroutine body)
(yield value)
(resume coroutine)

parser 分支与前面章节的特殊形式保持同一种写法。以 coroutine 为例:

1
2
3
4
5
6
if (head == "coroutine") {
    advance();
    auto body = parse_expr();
    expect(")", "expected ')' after coroutine body");
    return std::make_unique<CoroutineExpr>(std::move(body));
}

另外两项只改变节点类型与错误文字:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
if (head == "yield") {
    advance();
    auto value = parse_expr();
    expect(")", "expected ')' after yield value");
    return std::make_unique<YieldExpr>(std::move(value));
}

if (head == "resume") {
    advance();
    auto coroutine = parse_expr();
    expect(")", "expected ')' after resume operand");
    return std::make_unique<ResumeExpr>(std::move(coroutine));
}

这里故意不调用 validation。parser 只回答括号和节点形状是否成立;yield 所在的控制上下文要等整棵树构造完成以后才能判断。

三者都必须恰好有一个子表达式。少写或多写元素时,parser 会在右括号检查中报错。和 let、if 一样,这些纯字母拼写在列表 head 会被优先识别为特殊形式;在普通变量位置仍能作为名字出现,但不能借由重新绑定改变特殊形式的含义。coroutine_name_hygiene.lang 刻意绑定了这些名字,用来确认 lowering 生成的 $coro.* 不会与它们冲突。

对主例子运行:

1
2
cd code/11_coroutines
./mini ast examples/coroutine_protocol.lang

输出开头形如:

1
Let(g, Coroutine(Let(x, Yield(Int(40)), Add(Var(x), Int(2)))), ...)

这正好保留了读者关心的源码结构:x 的 initializer 是 yield,而两个 resume 发生在 coroutine 外部。

surface AST 不是解释器和后端的新输入语言

如果让 interpreter、自由变量分析、IR lowering 和汇编 emitter 各自处理 CoroutineExpr,很容易出现四种略有不同的暂停规则。本章只允许一个地方理解 surface coroutine:

1
2
3
4
5
6
7
surface AST
  ↓
validate_coroutines
  ↓
lower_coroutines
  ↓
chapter-ten core AST

mini ast 在这之前停止。其余命令都调用 lower_coroutines:

1
2
3
4
5
mini run     解释执行 core AST
mini core    打印 core AST
mini ir      把 core AST lower 成 IrProgram
mini alloc   分析同一份 IrProgram
mini compile 从同一份 IrProgram 生成汇编

为了防止实现者误绕开这条路径,解释器、free-variable analysis 和 IR lowering 一旦收到 surface coroutine kind,都会报内部错误。正常用户程序不应走到那里。

pass 的接口也把这条所有权边界写出来:

1
2
void validate_coroutines(const Expr& expr);
std::unique_ptr<Expr> lower_coroutines(std::unique_ptr<Expr> expr);

validation 只读 surface AST;lowering 接收并消费它的 unique_ptr,把子节点逐步搬进一棵新的 core AST。CLI 因此先处理 ast 命令,再执行 lower_coroutines(std::move(expr))。移动以后不能再使用原来的 surface tree,这与 unique_ptr 表达的独占所有权一致。

core 命令是本章的观察窗口

第十章用 alloc 把寄存器分配显示出来;本章增加:

1
./mini core examples/coroutine_protocol.lang

输出很长,这是正常的。开头会有类似:

1
2
3
4
5
Let($coro.state.0, Int(0),
  Let($coro.finish.1, Lambda(...),
    Begin(
      Set($coro.state.0, Lambda(...)),
      Lambda($coro.resumearg.7, ...))))

不需要逐字符背下这些名字。先抓住三件事:

  • $coro.state.0 是一个以后会被第八章 cell lowering 装箱的绑定;
  • Set($coro.state.0, Lambda(...)) 在安装下一步;
  • 用户写的 Resume(Var(g)) 已变成 Call(Var(g), Int(0))。

$coro. 含有 $ 和 .,而用户源码 identifier 只能是字母,所以没有用户变量能够意外捕获或遮蔽这些名字。这个性质叫作生成名字的 hygiene。

为什么不是 parser 直接生成 core AST

设想 parser 一读到 (yield 40),就立即生成一个 Set 和一个 Lambda。它会马上遇到两个问题:

  1. parser 还不知道这个 yield 是否在 coroutine 内、是否藏进普通 call;
  2. 生成“下一步”必须知道后面还有什么 let body、begin second 或 if branch,单独读到 yield 时信息不够。

因此顺序应当是:

1
2
3
parser                 保留源码结构
validation             检查 structured-yield 边界
coroutine lowering     看到完整上下文,构造 continuation

这也让错误信息更准确。例如 (yield 42) 先形成 Yield(Int(42)),随后才得到 yield used outside a coroutine,而不是被误报为一种奇怪的语法错误。

嵌套 coroutine 如何处理

lowering 遇到一个普通表达式中的 CoroutineExpr 时,会递归为它创建独立的 state、finish closure 和 resume closure。validation 则在进入内层 coroutine body 时重新开始检查 flow。

所以 nested coroutine 不需要全局的“当前 generator”变量;每个 coroutine 值都是一个普通 closure,并捕获属于自己的 state cell。这条设计直接复用了第七章的词法作用域规则:哪个 state 被使用,由 closure 的定义位置决定,而不是由当前调用位置猜测。


11.2 边界:哪些位置可以 yield

上一节

11.4 Lowering:把暂停后的计算装进 closure

下一节