重制 Minecraft 并非衡量标准
Recreating Minecraft Is Not a Benchmark

原始链接: https://kuber.studio/blog/Reflections/Recreating-Minecraft-is-Not-a-Benchmark

作者认为,人工智能的“演示基准”(即渲染弹跳球或 SVG 游戏控制器等病毒式传播的视觉任务)已不再是衡量模型真实能力的可靠指标。由于这些任务具有静态和可预测性,实验室可以轻易地通过对模型进行过拟合,使其在这些任务上表现出色,从而达到营销目的。这使得评估工作变成了一种预先安排好的“营销噱头”,而非对智能的真实测试。 尽管 GPQA 或开放排行榜等基准测试也面临数据泄露和“应试教育”等类似问题,但作者指出,病毒式传播的演示仍然是塑造公众认知的主要工具,因为它们能提供即时、易于理解的进度证明。 作者建议,真正的评估需要使用“保留”基准,即保持私密或定期更换问题的测试,以防止模型针对特定的已知问题进行优化。归根结底,虽然演示基准对营销有效,但作者敦促行业停止将病毒式的视觉成就与真实的模型智能混为一谈。如果一个模型在社交媒体上表现完美,却在实际工作流程中失败,那么这种演示不仅毫无意义,反而是一种干扰。

这场 Hacker News 讨论探讨了以特定任务(如重现《我的世界》)作为人工智能模型性能基准的局限性。 贡献者们的共识是,公开基准测试正变得日益不可靠,因为实验室可以轻易地通过优化模型来“刷分”。然而,用户认为像 GPT-6 Astra 这样的现代模型展现出了真正的通用能力,能够成功处理从定制游戏开发到根据照片生成复杂 3D 资产等各种创意任务。 一个重要的结论是,公开基准测试已在很大程度上沦为“内容营销工具”。随着模型能力的趋同,行业专业人士正逐渐摒弃标准化测试,转而依赖针对自身特定需求定制的私有评估集。归根结底,虽然作为公开指标的“基准测试”正失去公信力,但底层技术在通用的实际应用中已证明了其强大的能力。
相关文章

原文

GPT Astra released a couple of days ago and, inevitably, within the hour my entire feed was the same five things: recreating Minecraft in one prompt, painting themselves in MS Paint, the pelican riding a bicycle as an SVG, a ball bouncing in a rotating box with believable gravity, and an SVG game controller.

On paper these look like harder, more visual problems for a model to solve, there’s a reason they’re as big as they are. I’ve started calling them demo-benchmarks, visual and understandable enough for everyone to get but finite enough for the next model to be “perfect” on.

That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.

It’s not really their fault either, honestly I’d say it’s dumb if they didn’t - nothing sells a launch like a pelican or a 3D game controller the timeline can’t stop quoting.

A fixed, famous target and eight weeks of runway is a solved pelican, these tests never change and anything that never changes can be overfit. Every launch cycle proves it again.

A test you can perfect on a schedule measures preparation instead of capability, to me that’s anti the very definition of a benchmark, it should be a hard test, something very hard to perfect.

Claude Fable 5 recreating Minecraft in one prompt

The same dynamic runs through the open evals, smaller models that feel dumber in practice still outscore better ones on sites like Artificial Analysis. This isn’t hypothetical, Thinking Machines’ Inkling Small scored within a point of its flagship sibling on the Artificial Analysis Intelligence Index with less than a third of the parameters and beat it on Humanity’s Last Exam, GPQA Diamond and SciCode. Public, static, famous test sets leak into training data and fine-tuning choices.

Inkling Small, a third of the size, matching its flagship on the Artificial Analysis Intelligence Index

A launch is a first impression and first impressions are marketing, that’s why you’re always bound to be shocked - the shock was scheduled.

So what’s the alternative? Honestly, I’m not sure

because if you think about it, the obvious fix somewhat already exists. LiveBench rotates its questions, ARC-AGI keeps a private set, Humanity’s Last Exam holds part of itself back. Tests where the tested party doesn’t know what’s being tested: you can’t teach to a test that hasn’t been written yet.

But if holdout evals are the answer, though, why does the pelican still win?

The answer is because most of social media doesn’t need to understand a research paper to notice that the bicycle finally has pedals or is animated. A demo benchmark makes it obvious why it’s a better capability in seconds.

So no, there is no alternative to a good demo. But there are better alternatives to demo-benchmarks for seeing how good a model actually is. And if a model scores well but keeps failing at your work, that is actually a gap that deserves investigation.

Demo-benchmarks make great content. I just wish we’d stop grading with them.

联系我们 contact @ memedata.com