如何在 2026 年 9 月加速 Rust 编译器
How to speed up the Rust compiler in September 2026

原始链接: https://nnethercote.github.io/2026/09/30/how-to-speed-up-the-rust-compiler-in-september-2026.html

2026年7月29日至9月28日,Rust 编译器的性能有了显著提升:平均墙钟时间下降了4.57%,629项基准测试中有555项得到改善,部分测试的性能更是提升了两位数。 主要改进包括 rustdoc 优化、Clippy 的配置文件引导优化(PGO,速度最高提升18%),以及升级到 LLVM 23(平均耗时减少1.2%)。尽管 Nightly 新引入的 Polonius Alpha 借用检查器和 trait 求解器导致部分测试变慢,其中 `serde` 尤为明显,但后续优化已使 `serde` 的指令数减少了3%至5%,并显著改善了某些 trait 求解器表现异常的测试。 其他值得注意的性能提升还来自特化图和增量加载优化、分配次数减少,以及一种新的数据流遍历算法。该算法将一个包含18,000个基本块的函数所需的借用检查工作量削减了约94%,使检查速度提高了约30%。增大编译器栈大小以及一些细微的 AST/HIR 清理也带来了帮助。 即使是性能改进相关的 PR,其新增速度也开始超过 CI 逐个合并它们的速度,因此出现了汇总 PR。作者还宣布,接下来将负责 Rust 编译器性能优化工作。

The Hacker News 上的讨论涉及文章《如何在 2026 年 9 月加速 Rust 编译器》。 commenters 询问 Rust 编译为何较慢,尤其是与 C 语言相比时;同时也询问哪些性能瓶颈占主导地位。但摘录中没有提供明确的比较。 其中最有实质内容的评论赞扬了一项优化:通过改变编译器遍历控制流图的方式,将 `apply_effects_in_block` 的调用次数从约 150 万次降至 9 万次。这表明,替换低效算法或消除不必要的工作,效果可能胜针对少量热点循环进行微调。另一位评论者对按需 Polonius 和 trait solver 分析表示欢迎,并指出,在两个月的 629 个基准测试中,平均编译时间缩短了 4.57%。 总体而言,讨论表明,增量编译仍能带来较容易获得的性能提升,但未来的改进可能需要更加根本的算法变化。
相关文章

原文

My last post on the Rust compiler’s performance was two months ago and a lot has happened since then.

Overall progress

The measurements for the period 2026-07-29 to 2026-09-28 can be seen here.

The mean wall-time reduction was 4.57%, which is a remarkable improvement in just two months. Of the 629 benchmark measurements, 555 of them improved and only 74 regressed. A number of benchmarks saw double-digit percentage reductions. The technical term for this result is “a sea of green”.

rustdoc

In my last post I mentioned how Noah Lev got some enormous speed wins on rustdoc. He recently wrote a post explaining in some detail exactly how he did this. It’s an interesting and satisfying read.

Clippy

#159642: In this PR Jakub Beránek enabled PGO for Clippy, giving wall-time improvements across most Clippy benchmarks, in the best case by 18%!

LLVM update

#158734: In this PR Nikita Popov upgraded the LLVM version used by the compiler to LLVM 23. As often happens when we upgrade LLVM, we saw some nice speedups. The mean wall-time reduction across all benchmarks was 1.2%, which might not sound like much but is really impressive for a single PR. Great work from the LLVM folks!

The new borrow checker

The new borrow checker, Polonius Alpha (no relation to Napoleon Dynamite), was enabled on Nightly. It is more precise than the existing borrow checker and accepts some valid programs that the old borrow checker would reject. It does do more work than the old borrow checker, enough to make a measurable difference to compile time in a minority of cases, including the popular serde crate. Fortunately, Jack Huey has been on the case.

#161938: In this PR Jack made some liveness computations lazy, which reduced instruction counts for serde by 3-5%, and for some other benchmarks by less than 1%.

#163027: In this PR Jack adjusted a data structure and tweaked some inlining, for mostly sub-1% instruction count reductions across numerous benchmarks.

There is more work to be done to reduce the remaining Polonius Alpha regressions, but it’s worth noting that the “sea of green” shows these regressions were swamped by the many other recent improvements.

The new trait solver

The new trait solver, Penelope Hammertime, [Ed. note: is that right?] was also enabled on Nightly.

As I said, a lot has been happening.

Like the new borrow checker, the new trait solver is slower in a minority of cases. Jana Dönszelmann wrote a detailed post about the efforts to improve the performance of this new solver.

Jana’s post is detailed enough that I won’t say much more about the large amount of ongoing work on the new solver, but I will mention in passing the PRs I made: #160479, #160605, #160801, #160892, #161077, and #161211. Some of these reduced compile times greatly for certain outlier crates: 50% here, 25% there, 15% there, and even more on one stress test. And I am not the only one who has made progress here… go read Jana’s post.

xmakro

New contributor xmakro continued their run of good improvements.

#157281: In this PR xmakro optimized impl handling when building the specialization graph. This gave a mean cycle count reduction of 1.58% across all benchmarks, which is huge for a single PR.

#158059: In this PR xmakro optimized one aspect of the loading of incremental compilation data, reducing instruction counts across multiple benchmarks, in the best case by 6%.

#160473: In this PR xmakro avoided some allocations in a hot obligations processing path, reducing instruction counts across numerous benchmarks, in the best case by 2%.

#160268: In this PR xmakro avoided a lot of allocations by changing the old/new trait solver selection code to use static dispatch instead of dynamic dispatch. This gave mostly sub-1% instruction count reductions across a number of benchmarks. This hot allocation path had been showing up in profiles for a while and I had earlier tried exactly the same idea in #155714. But I got regressions on a couple of benchmarks, possibly due to slightly different choices of where to place some #[inline] attributes. It was good to see this obvious inefficiency fixed.

Dataflow analysis

#160193: In this PR I changed the CFG traversal algorithm used by the dataflow analyses in the compiler. These analyses iterate to a fixpoint and the traversal algorithm can affect how quickly the fixpoint is reached. For most code the new algorithm makes no difference, but the cranelift-codegen crate has one enormous function with over 18,000 basic blocks. The old algorithm required 1.5 million calls to apply_effects_in_block to reach a fixpoint for the EverInitializedPlaces analysis used by the borrow checker; the new algorithm requires 90,000. This gave an enormous ~30% wall-time reduction for a check build of this crate.

#160033: In this PR I made EverInitializedPlaces more efficient again, this time by not tracking unnecessary data for projections. This reduced instruction counts on the match-stress benchmark by 17%, and on a few other benchmarks by less than 1%.

LLMs

They’ve gotten very good at certain kinds of analysis. I’m still writing all my own code and text, because (a) that’s paramount, and (b) the project policy requires it, but I had useful LLM analysis assistance on several of the PRs mentioned in this post.

Anyway, enough about that.

Miscellaneous

#160535: In this PR Chris Denton increased the default stack size used by the compiler, which allowed the removal of ensure_sufficient_stack, a manual stack extension mechanism sprinkled about in places prone to high levels of recursion. There was a lot of discussion about this one because it can be difficult to decide how to best deal with stack exhaustion. But the performance effects are clear, with reduced instruction counts across many benchmarks, in the best case by almost 3%.

#160506: The project uses a lot of “rollup” PRs, where multiple PRs are merged together. This is because we don’t have sufficient CI capacity to merge every PR individually. Normally PRs that affect performance are merged by themselves so we can measure their effects clearly. For the first time ever, at one point we had so many performance improvement PRs waiting in the merge queue that Jonathan Brouwer created a rollup containing 10 performance-improving PRs to keep things moving! This is a good problem to have. And later on we had #162859 which contained four performance-improving PRs. (You needn’t worry about unexpected effects slipping in because we have the ability to run the perf benchmark suite on the individual PRs after merging, to make sure each PR had the expected performance effect.)

#162747: In this PR I made some minor improvements to the code that lowers AST to HIR. It was a cleanup that wasn’t expected to affect performance but it reduced instruction counts across numerous benchmarks, in the best case by 1.5%. Sometimes you get lucky.

Job status

Tomorrow I will start working at Hexcat on the compiler performance optimizations project goal. It’s exciting! Many thanks to Mara Bos, Predrag Gruevski, and all the other people who helped make this happen.

Editor’s postscript

The new solver’s name is not Penelope Hammertime; that was a joke.

Author’s postscript

Its real name is Pineapple Häagen-Dazs.

联系我们 contact @ memedata.com