Inside the Lume C11 Compiler: From a Tree-Walking Interpreter to Dual LLVM Backends

🇨🇳 中文版

Lume-core is the host-independent upstream tree of the Lume language: it factors the language itself out into its own tree — the frontend (lexer / parser / typecheck), a tree-walking interpreter, and two native code emitters (hand-written LLVM IR text + the libLLVM C API). It links only libc (and additionally libLLVM at build time if llvm-config is found), with no HTTP service, no agent runtime, and no Docker image. It stands in a "same language, different deployment shape" relationship with the host tree lume (the distribution that embeds the language into agent-httpd): language changes land here first, and the host tree then syncs from here.

Below, we walk through the real source code (src/, about 38 .c files, roughly 22k lines of C11 measured) layer by layer along the data flow.

TL;DR

  • Three layers: Frontend (lexer / parser / typecheck) → tree-walking interpreter (interp.c) → backends (hand-written IR text codegen*.c + libLLVM C API llvm_codegen.c).
  • Two backends deliberately coexist and cross-validate: --compile goes through libLLVM (LLVMVerifyModule checks the IR shape at build time), --compile-text goes through hand-written IR text handed to clang; make native-bench compares the two with the same script.
  • --no-pass skips the LLVM optimization pipeline, making it easy to observe the lowering produced by the emitter.
  • Frontend files shared with the host tree (hand-synced, not automated): lexer.c / token.c / parser_stmt.c / typecheck_stmt.c are still byte-identical (only vdom.c differs); both trees split the parser as parser.c + parser_expr.c + parser_stmt.c, with core's parser.c being the 696-line main file.
  • int is a real i64 (two's complement); the lexer routes integers through strtoll, and only literals with . / e / E go through strtod to become float — a deliberate design to avoid silent rounding above 2^53 and divergence between the interpreter and the two backends.

Overall Architecture

Lume's source is organized in three layers: frontend → interpreter → backend. The README states clearly that this tree is host-independent, linking only libc (with libLLVM added when llvm-config is found), carrying no HTTP service or agent runtime. Compared with the host tree (work/lume/lume), the frontend files here are hand-maintained copies, not automatically tracked forks — and they have already drifted: lexer.c / token.c / parser_stmt.c / typecheck_stmt.c remain byte-identical (only vdom.c differs), while interp.c / value.c / typecheck.c / loader.c and others have diverged.

The frontend's abstract data structures are all declared in lume.h, including the core types VM, Node, and Type, which are shared by the lexer, the recursive-descent parser, the type checker, the interpreter, and both backends, forming a unified cross-layer intermediate representation. The symbol table is maintained in loader.c and value.c: the type-checking phase fills in binding information, and the interpreter queries that same symbol table at runtime.

Frontend Tour

Lexer (lexer.c) The lexical phase converts the source program into a token stream. One noteworthy real design is the split between integers and floats — int is the language's true i64 and must never pass through double (strtod silently rounds above 2^53, which once caused the native backends and the interpreter to disagree on the same literal). Only literals with . / e / E go through strtod to become float:

bool is_float = false;
for (int i = 0; i < len; i++)
    if (buf[i] == '.' || buf[i] == 'e' || buf[i] == 'E') { is_float = true; break; }

Token *t = &o->toks[o->count - 1];
t->is_int = is_float ? 0 : 1;

Integers outside the i64 range error out directly rather than silently wrapping — because a value a source program cannot represent is a source bug, and silent mutation was precisely the root cause of the two backends diverging in the first place.

Recursive-descent parser (parser.c, parser_expr.c, parser_stmt.c) This tree splits parsing into parser.c (the 696-line main file) + parser_expr.c + parser_stmt.c, internally handling expressions and statements through recursive descent, building the Node AST defined in lume.h. The host tree uses the same split (parser.c + parser_expr.c + parser_stmt.c + parser_internal.h), which is exactly why the "hand sync" of the two trees' frontends drifts most easily across these files.

Type checker (typecheck.c, typecheck_expr.c, typecheck_stmt.c) After parsing, the type checker walks the AST, performs static inference and consistency checks against the Type system in lume.h, and attaches each node's inferred type to the AST, providing type information for backend code generation.

Middle-end Tour: The Tree-Walking Interpreter (interp.c)

The interpreter directly executes the AST using the classic tree-walking approach. eval_expr() dispatches by node type with a large switch — the very skeleton of "tree walking", where each case corresponds to a syntactic node:

static void eval_expr(VM *vm, Node *n, Env *env) {
    if (vm->error) return;
    switch (n->type) {
        case N_LITERAL:
            eval_expr_literal(vm, n);
            return;
        case N_VAR: {
            int found = 0;
            Value v = env_get(env, n->as.var.name, &found);
            if (!found) {
                vm_set_error(vm, "line %zu: undefined variable '%s'", n->line, n->as.var.name);
                return;
            }
            vm_push(vm, v);
            return;
        }
        case N_ASSIGN: {
            eval_expr(vm, n->as.assign.value, env);
            if (vm->error) return;
            Value v = vm_peek(vm, 0); /* keep rooted on the stack */
            env_set(vm, env, n->as.assign.name, v);
            return; /* result: assigned value, already on stack */
        }
        /* …N_ASSIGN_MEMBER / N_BINARY / N_CALL / N_IF / N_WHILE / N_FOR … */
    }
}

Execution depends only on the interpreter's own runtime stack (vm_push / vm_peek), generating no intermediate code; on return it pops the stack frame and passes the return value to the caller. This stack-based execution is also the semantic baseline the native backends must align with.

Backend One: Hand-Written LLVM IR Text Emission (codegen.c and the codegen_*.c family)

One backend directly concatenates the text form of LLVM IR to produce the intermediate representation. codegen.c is just the entry point — the real emitters are split by responsibility into adjacent files, sharing the same CG context through codegen_internal.h:

/* codegen.c — the text backend's entry point.
 *   codegen_types.c   Type -> LLVM spelling
 *   codegen_expr.c    every expression form
 *   codegen_scan.c    static types of expressions, block scanning
 *   codegen_sig.c     signature inference
 *   codegen_stmt.c    statements, loops, function bodies
 */

It walks the typed AST and emits the corresponding LLVM instruction sequence for each expression or statement, writing it into a .ll file, which is then handed to clang -S / llvm-as for assembly and linking. Because it is pure text concatenation, this path is easy to inspect by hand to verify the generated IR.

Backend Two: libLLVM C API Emission (llvm_codegen.c, backend_llvm.c)

The other backend uses LLVM's C interface directly to build IR in memory. The context struct in llvm_codegen.c holds the module, builder, and declared types:

typedef struct {
    LLVMContextRef    ctx;
    LLVMModuleRef     mod;
    LLVMBuilderRef    ab;          /* the moving builder                     */
    LLVMValueRef      fn;          /* function whose body is being emitted   */
    LLVMBasicBlockRef cur;         /* block the builder points into          */
    LLVMTypeRef       i1, i64, dbl, i8ptr, void_ty, i32;
    /* …locals / gvars / sigs … */
} Cg;

Each AST node calls the corresponding LLVMBuild* API to construct values (LLVMValueRef) and basic blocks. For example, emitting binary addition is a single, plain builder call:

case OP_ADD: r = LLVMBuildAdd(g->ab, a.v, b.v, "a"); break;

The IR built this way lives in LLVM's in-memory data structures, skipping the text-parsing step. backend_llvm.c handles backend initialization, target-machine selection, and hands the generated module to LLVMVerifyModule for on-the-spot validation at the point of production — a malformed IR shape is rejected as soon as it is produced, rather than surfacing only at the clang step.

Dual-Backend Coordination and Cross-Validation

Both backends share the same typed AST produced by the frontend, so the LLVM IR they generate should be semantically equivalent. As docs/NATIVE.md explains, the project deliberately keeps both backends to cross-validate each other: you can enable both backends on the same script to generate IR and diff them, catching any deviation in either implementation (tests/native_backends.sh does the interpreter / text / libLLVM three-way comparison, and scripts/check-backend-parity.sh checks whether the two emitters' AST-label coverage is consistent).

An updated conclusion (2026-10-03): the libLLVM path now runs the optimization pipeline — the entry point for LLVMRunPasses is in llvm-c/Transforms/PassBuilder.h, and backend_llvm.c already includes it, so it is no longer a matter of "calling into a crash inside a pure-C translation unit". Empirically, the text path and the libLLVM path land in the same tier on the run side (default<O2> / clang -O2), and the libLLVM path actually compiles faster; make native-bench gives the following measurements on the same nested loop (1e8 inner iterations):

backend   compile-ms   ir-KiB  obj-KiB  bin-KiB    run-ms  nopass-ms
text            446        3        1       35         9          -
llvm            184        2        1       35         9        215

Once the run side is even, the real sentinel is the last column nopass-ms: after --no-pass (equivalently the env var LUME_NO_PASS=1) skips the pipeline, the libLLVM leg jumps from 9ms to 215ms — if the pipeline silently fails one day, this column will jump first. Do not quote the old numbers (e.g. 220ms / 7ms) directly: those were the table before the pipeline was connected, and the docs are explicitly marked as outdated. You can still switch to --compile-text when throughput is sensitive, but the reason now is no longer "libLLVM is slow because it doesn't run a pass".

Compilation Flow in Reverse (main.c)

From the CLI entry main.c, when the user runs lume --compile script.lume, the program first reads the file, runs lexing / parsing / type checking to obtain the typed AST, then enters the corresponding code generation based on the selected backend. main.c explicitly models "which native backend" with an enum — the default is libLLVM, falling back to text when built without libLLVM:

/* Native backends the CLI can ask for. --compile picks the default, which is
 * the libLLVM one when the binary was built with libLLVM (LLVM then verifies
 * the IR shape while it is built); --compile-text is the opt-out. */
enum { NAT_DEFAULT = 0, NAT_LLVM = 1, NAT_TEXT = 2 };

The flags exposed by the CLI are consistent with the README: --check (type check only), --dump (print the IR produced by the text backend), --compile / --compile-llvm / --compile-text, --no-pass (skip the optimization pipeline), --no-fs / --no-net (narrow the script's capability boundary). main.c then temporarily invokes the system's cc (or libLLVM's JIT) to turn the IR into an executable binary and return it to the user.

Build, Test, and Platform Support

Dependencies are minimal: a C11 compiler (cc), and optionally llvm-config (determining whether --compile-llvm is available). Common targets:

make             # build bin/lume-core (needs cc; links libLLVM only if llvm-config present)
make check       # type-check the built-in examples, produces no artifacts
make test        # parity + unit tests + crypto + both native backends + consistency
make asan        # rebuild with ASan/UBSan and run unit tests
make native-bench # compare the two backends on the same script

Windows (MSYS2 / mingw-w64) is fully supported: the interpreter, type checker, both native backends, and module import all work under mingw. File/directory/lock/time builtins (mkdir / file read-write / files / lock_file / strftime) are also available, but they go through shims in os_win32.c (lume_mkdir / lume_flock) rather than builtins_fs.c writing the Win API directly. The only unsupported pieces are --watch (which depends on fork/exec/kqueue) and outbound HTTP: specifically http_get / http_post / … — src/builtins_http.c is excluded at build time and compiled with LUME_HAS_HTTP=0, so these builtins pass type checking in such a build but report a clear error when called (rather than failing to link outright).

Source Navigation

Layer Files
Core types lume.h
Frontend lexer.c · token.c · parser.c · parser_expr.c · parser_stmt.c · parser_internal.h · typecheck.c · typecheck_expr.c · typecheck_stmt.c · vdom.c
Interpreter interp.c · value.c · loader.c · rt.c
Text backend codegen.c · codegen_types.c · codegen_expr.c · codegen_stmt.c · codegen_scan.c · codegen_sig.c · irbuf.c · backend.c
libLLVM backend llvm_codegen.c · backend_llvm.c · backend_llvm.h
Builtins builtins.c · builtins_str.c · builtins_math.c · builtins_fs.c · builtins_http.c · builtins_crypt.c · builtins_hof.c · builtins_catalog.c
Bridge / entry bridge_stub.c · bridge_native.c · bridge_mcp.c · bridge_lsp.c · bridge_serve.c · main.c
Docs docs/LUME.md (user guide) · docs/DEVELOPMENT.md · docs/ARCHITECTURE.md · docs/NATIVE.md · docs/SPEC.md · docs/STYLE.md · docs/PITFALLS.md

FAQ

Does this Lume tree depend on any library other than libc?

It additionally links libLLVM only when llvm-config is detected; otherwise it depends only on the standard C library. No Node, no server runtime, no SQLite.

What is the main difference between this tree's parser and the host tree's parser?

Both trees split the parser as parser.c + parser_expr.c + parser_stmt.c (core’s parser.c is the 696-line main file); the two trees’ frontends are hand-synced copies, not automatically tracked, so files like parser_stmt.c must be merged manually as core changes.

How does the interpreter handle function calls and returns?

The interpreter uses the runtime stack (vm_push / vm_peek) to hold return addresses and local variable frames, and on return it pops the stack frame and passes the return value to the caller.

What is the main difference between the two backends?

The text IR backend generates a .ll file by directly concatenating LLVM IR strings and then hands it to clang; the libLLVM backend calls the LLVM C API to build LLVMValueRef and basic blocks in memory, validating on the spot with LLVMVerifyModule at the point of production.

How can I verify that the IR generated by the two backends is equivalent?

You can enable both backends on the same script to generate IR and then diff them; tests/native_backends.sh does the interpreter / text / libLLVM three-way comparison, and scripts/check-backend-parity.sh checks whether the two emitters’ AST-label coverage is consistent.

What does the –no-pass option do?

It skips the LLVM optimization pipeline and outputs the raw, unoptimized IR directly, making it easy to inspect the lowering produced by code generation.

Why does the program built by the libLLVM backend run slower?

That was the old conclusion (before 2026-10-03). The libLLVM path now runs the optimization pipeline — the entry point for LLVMRunPasses is in llvm-c/Transforms/PassBuilder.h, and backend_llvm.c already includes it, so it is no longer a matter of “calling into a crash inside a pure-C translation unit”; on the run side the measurements are now even (default<O2> / clang -O2 in the same tier). The real sentinel today is the column when --no-pass (or LUME_NO_PASS=1) turns the pipeline off: the libLLVM leg jumps from roughly 9ms to roughly 215ms, and it will jump first if the pipeline silently fails one day. You can still switch to --compile-text when throughput is sensitive, but the reason is no longer “libLLVM is slow because it doesn’t run a pass” — it is the fallback path when the LLVM development package cannot be installed.

Project Links and Further Reading

Lume-core is the host-independent upstream tree of the Lume language, adjacent to the host distribution lume (current workspace path work/lume/lume-core); language changes land here first and are then synced by the host tree. By pipeline convention this preserves the remote address github.com/erishen/lume-core; if the project has not pushed a public remote, that link will 404 — this is a template convention, not invented by this article. The authoritative externally accessible entry point is the Lume series on erishen.cn (e.g. Lume: An Agent DSL Server in a Single C11 Binary).

Comments

Leave a reply

Your email address will not be published. Required fields are marked *

AI Engineering Practices & Open Source Projects

Shop Web Chat Nsbp About Privacy

@ 2026 ESN
沪ICP备2024079226号-1   沪公网安备31010502007082号