Show HN: TurboGPT:在 13 秒内训练 22KiB 的 Transformer 模型
Show HN: TurboGPT: train 22KiB transformer in 13s

原始链接: https://github.com/lostmsu/TurboGPT

Tiny byte-level GPT training in CUDA C++. MIT. Linux/NixOS: nix-build -o build/nix-result Windows, Visual Studio 2022 C++ tools, and CUDA 13.4: CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list. .\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4 The run stores its checkpoint at runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it. runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints. hn1g after 1.5G training tokens: 2.5295 BPB.

一个以 **TurboGPT** 为中心的 Hacker News 讨论项目。该项目受 minGPT 启发,据说只需 13 秒就能训练出一个仅 **22 KiB** 的 Transformer。它是一个字节预测器,可以在任意文件上进行训练,需要 **CUDA 13.4**,并且目前只提供了适用于 **Windows** 的构建脚本。 一些评论者质疑再做一个极简 GPT 实现的动机,认为 Karpathy 的 nanoGPT 或 minGPT 等项目已经具有更高的教育价值。另一些人则表示,这类项目可能是为了娱乐、给简历增添亮点,或只是因为人们喜欢反复实现熟悉的工具;还有人调侃,它很可能依赖 AI 生成的代码。 作者解释说,minGPT 适合用来学习 Transformer 架构,但在家里进行更大规模的实验需要更高的速度。TurboGPT 提供了一种快速测试架构想法的方法,尤其是当这些想法尚无经过优化的训练原语时;不过作者也承认,这种速度可能已经有些 overkill。 评论者还注意到,如今人们倾向于按照模型序列化后的文件大小,而不是参数数量来命名模型;有人甚至提出,在如此小的规模下,模型优化或许可以直接求解。
相关文章

原文

Tiny byte-level GPT training in CUDA C++. MIT.

Linux/NixOS:

nix-build -o build/nix-result

Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:

CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list.

.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4

The run stores its checkpoint at runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it.

runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.

  • hn1g after 1.5G training tokens: 2.5295 BPB.
联系我们 contact @ memedata.com