Turbo Haskell
Turbo Haskell

原始链接: https://comonad.com/reader/2026/turbo-haskell/

THC(“Turbo Haskell 编译器”)是一款实验性编译器和运行时,通过 Truffle 和 GraalVM 在 JVM 上执行经 GHC 优化的 Haskell Core。它支持全部 GHC 9.14.1 prim 操作、Template Haskell 和 Linear Haskell,并通过 GraalVM Native Image 提供即时编译和提前编译。它能够编译 Pandoc、Happy、Alex 和 GHC 本身等主要工具,也能处理 Cabal 包,包括多库包和 Backpack 包。 THC 通过 Sulong 为 Python、Ruby、R、JavaScript 和原生 C/C++ 提供多语言 FFI,并支持 `Data.Text` 与 Truffle 字符串之间零拷贝的 UTF-8 转换。其后端支持多线程、GHC 字节码、异步异常、异常屏蔽、`throwTo` 以及基于 Loom 的绿色线程。它还支持运行时宽度 SIMD 和尾调用优化,能够将热点递归转换为紧凑循环,同时定期压缩泄漏的栈帧。 早期预热后的基准测试结果显示,速度从比 GHC 快 3 倍到比 GHC 慢 3 倍不等,但后续的性能回退和非尾调用导致的栈增长仍是问题。开发地址为 `github.com/ekmett/thc`,相关讨论在 Libera Chat 的 `##thc` 频道进行。

The Hacker News 的讨论聚焦于 **Turbo Haskell(THC)**,这是 Edward Kmett 开发的一款实验性 Haskell-to-JVM 编译器。它复用了 GHC 的部分组件实现了 GHC 9.14.1 的原始操作,并为 GHC Core 加入了 JIT,可通过 JVM 运行 Haskell。该项目还强调与 GraalVM/Truffle 的多语言互操作性,包括零拷贝的 `Text`/`TruffleString` 转换,不过专用的 Java FFI 支持仍然有限。 评论者开玩笑说,这个名字让人联想到 Turbo Pascal、Turbo Vision,或者某款已被遗忘的蓝色 TUI IDE;也有人将其与 Frege 比较,并质疑是否还有必要再开发一款 JVM Haskell 编译器。另一场更广泛的讨论涉及有关代理在一周内生成数千次提交的报道,这引发了人们对代码质量和审核机制的质疑。Kmett 表示,代理完成了大部分初始实现和测试生成工作,但此后他将重心放在 CI、测试夹具、跨平台稳定性和代码清理上。他也承认,目前该项目的编译速度尚不具竞争力。项目配套的编辑器 `thc-edit` 提供了更传统的编辑体验。 其他讨论分支还涉及 Java 互操作性、AI 生成的 Haskell,以及为什么生态系统成熟度仍会影响人们为代理式编程选择编程语言。
相关文章

原文

Exactly a week ago (as a joke), I started writing THC, my “Turbo Haskell compiler,” while on vacation visiting Bartosz Milewski.

It has grown a tiny bit since then.

THC now implements every one of GHC 9.14.1’s prim-ops and provides a JIT for GHC Core that runs Haskell on the JVM. It uses the approach for running typed functional languages I developed several years ago in Cadenza (talk), using Truffle and GraalVM.

GHC still handles parsing, typechecking, desugaring, and Core optimization. THC takes over from there, compiling and executing that Core through its own runtime on Truffle/GraalVM. Advanced language features such as Template Haskell and Linear Haskell are fully supported.

While it can be used as a JIT for GHC-grade Haskell, it also supports ahead-of-time (AOT) compilation with Native Image, allowing it to produce executables.

THC is capable of JIT- or AOT-compiling a number of Haskell programs, including pandoc, happy, alex, and, as of today, even GHC itself!

THC resolves packages using Cabal and fully supports packages with multiple libraries, including Backpack.

Borrowing libraries

THC provides polyglot FFI to Python, Ruby, R, and JavaScript, letting Haskell raid libraries from other languages and bring their output straight into a JIT-compiled Haskell program. Conversion between Data.Text and Truffle strings over FFI is zero-copy for UTF-8-encoded strings inside other polyglot languages.

The idea is that if you need a data frame, want to run an LLM, or want a D3.js visualization, you should just pass Text out through foreign imports. Quasi-quotation-based inline-<language name> style bindings should be pretty easy to implement as well.

C/C++ bits in your Haskell libraries are run via FFI to native-mode Sulong (LLVM on the JVM). Managed-mode Sulong, where LLVM is interpreted inside the JVM and pointers are managed and garbage-collected, is also available, but is not used by the normal foreign import path.

Evaluation and concurrency

Internally, THC supports two different backends for Truffle evaluation: a bytecode-based JIT target and a traditional AST-based JIT target. Both can run in a single-threaded or multi-threaded style, with additional locking for the latter. It also supports GHC bytecode itself, so it can run BCO code as produced by GHCi.

THC fully supports throwTo, asynchronous exceptions that leave behind resumable code, and masking.

THC supports both “normal” Java threading and Project Loom, upon which it offers lightweight GHC-style green threading with a HEC-style runtime executor permitting cheap MVars and the like.

SIMD

THC supports SIMD using the relatively limited supply of available GHC prim-ops, but it can also go further, allowing runtime selection of the SIMD “species” width and JIT-compiling loops using that information through the incubating Vector API (jdk.incubator.vector). This lets the JIT start to earn its keep!

In fact, nothing prevents the runtime from providing complete RuntimeRep-polymorphic code at runtime other than the fact that we have no Core that takes advantage of that freedom!

Tail calls

Hot tail calls become loops. When execution has to fall back to ordinary calls, THC periodically unwinds the accumulated stack frames.

Previous efforts to run functional code on the JVM, such as the Eta programming language and the design I used with Runar Bjarnason for trampolining Scalaz’s monads, used a trampoline mechanism. THC instead uses a code transformation trick to keep hot tail calls inside tight basic-block-style loops with side exits.

During recursion in tracing mode, THC fills a 64-bit Bloom filter to detect likely recursive tail calls. When it finds a likely hit, it throws a slow-path exception to connect the continuation with the launch site, and then custom Truffle nodes get Graal to transform the current tail-call loop across function bodies into a tight loop. False positives mean extra slow-path work; they don’t change the program’s result.

When later code paths diverge, THC tries to grow additional side loops, like a tracing JIT, until it hits the JVM’s limits on function body size. At that point, it’ll finally spill a tail call in a way that “leaks” a stack frame.

That leak is temporary. We can compact the accumulated stack frames using a trick somewhat similar to CHICKEN Scheme’s garbage collection strategy, reusing the machinery we needed to support resumable code in the presence of asynchronous exceptions. The result is that hot loops can run very hot indeed.

When benchmarking Data.Map in particular, I found it needed something like 66 fallback trampoline calls compared to several million fast-path calls.

Performance

The runtime can also use compressed ordinary object pointers (compressed oops). These represent heap references as 32-bit offsets rather than full 64-bit pointers, reducing the memory occupied by references and helping more data fit in cache. With the JVM’s usual 8-byte object alignment, this limits the heap to roughly 32 GB when running in this mode.

Performance was a key consideration for the first couple of days of development. For tests on Data.Map and the like, I was able to get things to run within a general range of 3× faster to 3× slower after warmup, mostly hovering around 10–20% slower than GHC. That said, we haven’t been benchmarking for the last few days while we raced for broader coverage and suffered a 10× or so performance regression on some easy benchmarks in the meantime. Development effort continues to plug away at these to keep it under control.

We haven’t yet tested whether stack growth on non-tail-call paths remains bounded relative to GHC’s stack usage. Ensuring that bound remains possible future work.

Development

The code is available at github.com/ekmett/thc, with documentation covering how to build, run, and use THC.

Development is proceeding on irc.libera.chat in the ##thc channel.

Come join us.

—Edward Kmett

Discuss on Reddit.

联系我们 contact @ memedata.com