原文
Comprehensive harness evaluation
9 harnesses across 12 configurations, tested on identical software engineering tasks with the same model and runtime.
Identical cold start on every run
All 360 trials start from the same fresh checkpoint restore. Formal tasks were never run early, preventing warm-cache bias.
Neutral evaluation with no home-field advantage
Every run used Kimi K3 and was executed on Runta with a fresh restore using identical vCPU, memory, disk size, disk contents and memory state