Show HN:Yantra——一款用于 C++ 的 LALR(1) 解析器生成器
Show HN: Yantra – an LALR(1) parser generator for C++

原始链接: https://github.com/TantrixAuto/yantra

Yantra 是一款 C++23 编译器和 LALR(1) 解析器生成器,除标准库外不依赖任何其他库。它提供集成式推入式词法分析器、Unicode/UTF-8 支持、词法分析器模式、AST 构建以及多种自顶向下 AST 遍历器。它既可以生成包含 `main()` 的独立合并式解析器,也可以生成独立的 `.hpp` 和 `.cpp` 文件。 与 Bison、Yacc 和 Lemon 不同,这些工具在归约过程中自底向上执行语义动作;Yantra 会先构建完整的 AST,然后自顶向下遍历它,从而允许父节点的动作在子节点之前执行。它具备 LALR(1) 的效率,同时避免了 ANTLR 的 LL(*) 和自适应预测开销,并且仍然是原生 C++ 工具,不需要 JVM。不过,Yantra 较新、成熟度较低,仅支持 C++,并且并非为增量解析或 tree-sitter 一类的编辑器集成而设计。 使用 CMake 构建它以生成 `ycc`,然后从 `.y` 语法文件生成解析器,并使用 Clang、GCC 或 MSVC 编译。输入可以来自文件、字符串、交互式会话,或套接字等流式数据源。Yantra 采用 MIT 许可证,由 Renji Panicker 维护。

Yantra 是一个处于早期阶段的、采用 MIT 许可证的 C++23 LALR(1) 解析器生成器。它可以从一份语法规则生成词法分析器、解析器、抽象语法树(AST)以及 AST 遍历器。它与 Yacc、Bison 或 Lemon 的主要区别在于,会先构建完整的 AST,然后在独立的自顶向下遍历中执行语义操作。这样,父节点的操作可以先于子节点执行,并且支持多种遍历器,例如从同一次解析结果中生成 C++ 和 Java 代码。它还提供词法分析器模式、可选的合并输出,以及自动检测语法冲突的功能。 HN 上的讨论主要围绕为什么在生产级编译器通常更倾向于手写递归下降解析的情况下,仍然要使用生成的解析器。递归下降通常能提供更好的上下文相关错误处理、错误恢复能力和可调试性;而解析器生成器对于声明式 DSL 语法和多目标代码生成仍然很有用。评论者澄清/errors的质量更多地与自顶向下还是自底向上的解析方式有关,而不是取决于解析器是生成的还是手写的,并提到了 ANTLR 和 Chumsky 等工具。 Yantra 目前不支持错误恢复,生成的错误信息大多是通用的语法错误。其他建议包括添加 Python 绑定,以及修正文档中有关表达式结合性的说明。
相关文章

原文

CI License: MIT Version C++

Yantra Logo

Yantra is a powerful compiler compiler and LALR(1) parser generator written in C++, with the following core features:

  • An integrated lexer
  • Built-in support for UNICODE/UTF8 input
  • Built-in AST builder
  • Built-in AST walker(s)
  • Bottom-up parsing (being LALR), and top-down walking (traversal)
  • Multi-mode lexer, useful for implementing nested multi-line comments, etc.
  • Lexer-driven (push-based) parser: reads input one character at a time and feeds tokens to the parser as they complete, useful for processing input as it arrives (e.g. from a socket).
  • An optional amalgamated mode, where the entire parser is generated as a single cpp file, along with a full-featured main() function.
  • Or, in non-amalgamated mode, the parser is generated as separate .hpp and .cpp files, ready to drop into an existing project.

The name Yantra is Sanskrit for machine, as in state machine in this context.

Yantra has no dependencies beyond the C++ standard library, so building it is a plain CMake build:

git clone [email protected]:TantrixAuto/yantra.git
cd yantra
mkdir build && cd build
cmake ..
cmake --build .

This produces the ycc executable in bin/. Save a grammar file, hello.y:

start := stmts;
stmts := stmts stmt;
stmts := stmt;
stmt := ID;

ID := "[A-Za-z]+";
WS := "\s"!;

Then generate a parser from it:

bin/ycc -c ascii -f hello.y -a

This writes hello.cpp (an amalgamated, self-contained parser with its own main()) and hello.log. Compile it with any C++23 compiler:

# clang
clang++ --std=c++23 -o hello hello.cpp

# gcc
g++ --std=c++23 -o hello hello.cpp

# MSVC (cl.exe, from a Developer Command Prompt)
cl /std:c++23 /EHsc /nologo hello.cpp

The grammar above recognizes one or more whitespace-separated alphabetic words. -s <string> feeds that string directly to the parser as input (as opposed to -f <filename>, which reads from a file, or -i, which reads interactively from the console):

# succeeds silently
$ ./hello -s "hello world"
$ echo $?
0
# -t1 prints the parsed AST
$ ./hello -s "hello world" -t1
0:start_1(1:stmts_1(2:stmts_2(3:stmt_1(4:ID(hello))) 2:stmt_1(3:ID(world))) 1:_tEND())
# fails: ID only matches letters, "123" isn't valid input for this grammar
$ ./hello -s "hello 123"
s1-err:?a1.in(001,007):TOKEN_ERROR{{token: }}
hello 123
$ echo $?
1

Yantra parses the entire input into an AST first, then walks it top-down calling your semantic actions, unlike most parser generators, where actions run bottom-up as each rule is reduced. That ordering is what lets a parent rule's action run before its children are visited. Save this as calc.y:

%class Calculator;

start := expr;

expr := expr(a) PLUS expr(b)
%{
    std::cout << "Adding" << std::endl;
%}

expr := NUMBER(N)
%{
    std::cout << "Number: " << N.text << std::endl;
%}

NUMBER := "\d+";
PLUS := "\+";
WS := "\s+"!;

Generate and compile it the same way as above (bin/ycc -c ascii -f calc.y -a, then any of the three compiler commands), then run it:

$ ./calc -s "1 + 2 + 3"
Adding
Number: 1
Adding
Number: 2
Number: 3

1 + 2 + 3 parses left-associatively as (1 + 2) + 3. So the outer Adding, the root of the tree, prints first, followed by its left child (Number: 1) and then its right child, which is itself another Adding node with its own two children. A hand-written recursive-descent or bottom-up parser would have to build extra AST classes and a separate walking pass to get this ordering. Here it falls out of the grammar directly.

See the Build Instructions and Tutorial below for a real walk-through of the grammar syntax.

  • vs. Bison / Yacc / Lemon (the classic LALR(1) family):
    • Lemon, from SQLite, is Yantra's direct stated inspiration.
    • These run semantic actions during parsing, as each rule reduces, bottom-up.
    • Yantra always builds the full AST first, then walks it top-down in a separate pass, so a parent rule's action can run before its children are visited.
    • A single grammar can also define more than one walker (e.g. one that emits C++, another that emits Java, from the same parse).
    • Getting either of those out of the Bison family means hand-building your own AST and walker on top.
  • vs. ANTLR:
    • ANTLR walks a fully-built parse tree too, but that comes for free from its LL(*) algorithm, which already builds the tree top-down as it parses.
    • Yantra gets the same top-down walk out of LALR(1), a bottom-up algorithm with no natural "whole tree exists yet" moment during parsing, while keeping LALR(1)'s time and space efficiency over adaptive LL(*).
    • Beyond that, Yantra targets C++ only (ANTLR generates for many languages) and ships its own integrated lexer with mode-stack support instead of a separate lexer generator.
    • ANTLR's own generator tool is Java, so using it from a C++ project means adding a JVM to the build toolchain just to run the generator. Yantra is a native C++ executable with no such dependency.
    • ANTLR is far more mature and widely used. Yantra is a much smaller, newer, single-maintainer project.
  • vs. tree-sitter:
    • A different problem entirely. It's built for incremental, error-tolerant parsing embedded in editors and IDEs (what GitHub, Neovim, etc. use it for), not for generating a compiler/codegen backend.
    • Yantra doesn't do incremental reparsing and isn't trying to.

See Known Limitations for an honest list of what Yantra doesn't do yet.

The following are a set of key links to get familiar with Yantra.

It is recommended that they be read in the given order.

See https://github.com/TantrixAuto/lingo for standalone sample project that uses yantra.

Language Server Extension

This is a language server extension created by Raj Chaudhuri that provides syntax highlighting for Yantra files in vscode, qtcreator, and any other IDE that supports the Language Server Protocol.

https://github.com/rajware/yantra-language-server

Yantra is licensed under the MIT License.

Renji Panicker (@renjipanicker)

联系我们 contact @ memedata.com