一次由内存交换导致的 40 毫秒 Go 垃圾回收暂停。
A 40ms Go garbage collector pause caused by swap

原始链接: https://frn.sh/go-gc/

在 Go 工作负载运行时,使用交换空间(swap)吸收内存峰值可能会造成明显干扰。Go 的垃圾回收器(GC)会在停止世界(stop-the-world)阶段反复读取持久化的运行时元数据。由于这些页面会被重复使用而不是释放,内核的页面老化机制可能会将它们换出到交换空间,从而迫使 GC 在所有 goroutine 都暂停期间处理同步的严重页面故障。 在配备 MGLRU 和 NVMe 的 Linux 6.8 上,GC 暂停时间的中位数约为 51 微秒,但最长暂停时间达到 40 毫秒。其中,39 毫秒用于处理 GC 内部记账期间发生的 228 次页面故障。此类暂停会显著增加尾部延迟,并推迟已完成 I/O 的处理。 内存分配密集的路径也受到了影响:构造一条 511 KiB 消息所需的时间,在 NVMe 上从 3–5 毫秒增加到 105 毫秒,在网络存储上则增加到 903 毫秒,不过这一额外开销仅影响执行内存分配的 goroutine。 结论是,交换空间本身并不一定有害,但将 Go GC 元数据换出到交换空间,可能会把存储延迟转化为全局停止世界延迟。在此次实验中,Go 1.26 的 Green Tea 垃圾回收器几乎没有带来改善。

Hacker News 最新 | 往期 | 评论 | 提问 | 展示 | 工作机会 | 提交 登录 一次由交换导致的 40 毫秒 Go 垃圾回收器暂停 (frn.sh) 10 分 由 shellpipe 发布 2 小时前 | 隐藏 | 往期 | 收藏 | 讨论 | 帮助 考虑申请 YC 2027 年冬季批次! 申请截止至 11 月 2 日。 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请加入 YC | 联系我们 搜索:
相关文章

原文

I’m writing this so you don’t slap your forehead like I almost did when I decided to run swap in production to absorb memory spikes.

I had a cgroup with two processes: one is a Go process that calls io.ReadAll and then proto.Unmarshal, creating a blob and then a graph struct (which is marked as scan by Go’s allocator). The other process is an HTTP server that mostly stays quiet.

Whenever the collector runs, it reads those scan spans, pointer by pointer, and decides what to do with them. So I thought: ok, under memory pressure, the kernel is going to evict pages to the swap device, but since the eviction is per cgroup, and not per process, both processes’ pages are going to be evicted - so there is only a small chance that this will turn into a sad dance of swap-in and swap-out between the kernel and the garbage collector.

I was wrong. While experimenting with this, I found a problem that could’ve hurt me: Go’s garbage collector reads its metadata (outside the heap, in a region that is not freed) in a stop-the-world pause, and that metadata can be in swap.

I did a mock run on a Hetzner box using kernel 6.8 with MGLRU enabled1 You can find everything about these experiments: plots, the mock allocator, bpf scripts, python scripts, etc., here: https://github.com/frnsimoes/go-gc-swap-cost. The median pause was around 51 us. With the metadata on the NVMe, the worst pause was 40ms.

alt

To check where those 40ms went, I wrote a small bpf script that counts page faults while the world is stopped. This was the worst one: 39902 us, faults during it 228, 39013 us in faults. 39 of those 40ms were spent in 228 page faults. Those faults happened inside the GC’s bookkeeping:

// addr2line output
0x42e5c8  runtime.(*spanSet).reset         /usr/local/go/src/internal/runtime/atomic/types.go:194
0x4219de  runtime.finishsweep_m            /usr/local/go/src/runtime/mcentral.go:71
0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724
0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518
0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722

0x4276a4  runtime.nextMarkBitArenaEpoch    /usr/local/go/src/runtime/mheap.go:2481
0x421a65  runtime.finishsweep_m            /usr/local/go/src/runtime/mgcsweep.go:268
0x4629cf  runtime.gcStart.func2            /usr/local/go/src/runtime/mgc.go:724
0x46dd8a  runtime.systemstack              /usr/local/go/src/runtime/asm_amd64.s:518
0x4169dc  runtime.gcStart                  /usr/local/go/src/runtime/mgc.go:722

That’s a potential failure mode. Go’s GC has to stop the world at two points: when it performs a sweep termination, and when it performs a mark termination. We had 312 of those pauses in 30 minutes.

So here is why this happens: the runtime allocates those pages. They are not freed, but reused. Those pages are read in GC cycles. Because the kernel evicts pages by age, it sends the least recently accessed pages to swap. The GC runs, stops the world, tries to read those pages, but now we have a major page fault. The kernel needs to read PTEs, and then call do_swap_page, find a new frame, charge it to the cgroup, read the pages, submit a bio, wait for the disk, and put them back in memory - just to keep it short.

Those 40ms seem harmless at first. But we are talking about a stop-the-world pause. Those 40ms mean everything has stopped - in Go’s terminology, every P has stopped, so, for example, if a goroutine was waiting for I/O, during that pause the I/O might return and there would be no one to handle it. 40ms is 800 times the median pause. It happens two or three times per memory spike during the test. It is a lot.

And then I noticed another thing: building one 511 KiB message, which usually takes 3-5 ms, jumped to 105 ms on the NVMe and 903 ms on Hetzner’s network volume. Per message, this costs more than the metadata pause. But only the goroutine doing the allocation pays that price, so at least it’s localized, and not global like the metadata one.

alt

I haven’t confirmed where that time goes, but even so I wanted to mention it here, because that’s another cost you would have to pay. In any case, I agree with Chris Down, swap is not evil. But, yeah, it didn’t behave well with garbage collection, and, in production, I’m collecting a lot.


Update, September 14.

Someone asked me if Go 1.26’s Green Tea garbage collector changed the way the GC reads metadata. I measured it, and the impact is negligible.

alt
联系我们 contact @ memedata.com