Show HN: FrontierHarness Eval – 9 种评测方案,同一模型,单次成本差异达 17 倍
Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

原始链接: https://frontierharness.org

综合基准评估 在 12 种配置下对 9 种工具集进行评估,均使用相同的软件工程任务、模型和运行环境。 每次运行均采用相同的冷启动 全部 360 次试验均从全新的检查点恢复启动。正式任务从未提前运行,排除了缓存预热带来的偏差。 无主场优势的中立评估 每次运行均使用 Kimi K3,并在 Runta 上执行,通过全新的恢复方式,确保 vCPU、内存、磁盘大小、磁盘内容及内存状态完全一致。

相关文章

原文

Comprehensive harness evaluation

9 harnesses across 12 configurations, tested on identical software engineering tasks with the same model and runtime.

Identical cold start on every run

All 360 trials start from the same fresh checkpoint restore. Formal tasks were never run early, preventing warm-cache bias.

Neutral evaluation with no home-field advantage

Every run used Kimi K3 and was executed on Runta with a fresh restore using identical vCPU, memory, disk size, disk contents and memory state

联系我们 contact @ memedata.com