Cognition 的 SWE-2 在 Terminal-Bench 2.1 上取得了 92.8 分。
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

原始链接: https://tokenstead.ai/models/swe-2

Cognition 推出了 **SWE-2**,这是一个基于 Kimi K3 底层构建的 2.8 万亿参数混合专家模型(MoE)。通过将强化学习(RL)扩展至万亿参数规模,Cognition 在大多数基准测试中将基准性能提升了 5–6 个点。 该模型采用了先进的推理技术,包括 FP8 内核和预填充延迟器(prefill delayer),以提高吞吐量。值得注意的是,SWE-2 在代理编码任务中实现了高效表现:它超越了之前的 SWE-1.7 版本,同时交互轮次减少了 58%,成本降低了 81%。Cognition 声称其在 FrontierCode 上的表现具有竞争力(50.0),仅以微弱差距落后于 Claude Fable 5.1 和 GPT-6 Astra 等顶级模型,且运行成本显著降低。 然而,SWE-2 在长视程任务上表现吃力,其在 Terminal-Bench 4.0 上的得分低于前沿模型便证明了这一点。该模型目前为闭源,仅可通过 Devin 生态系统(桌面端、CLI 和 Fusion)使用,暂不提供公共 API。所有性能数据均由 Cognition 自行报告,尚待独立验证。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 Cognition 的 SWE-2 在 Terminal-Bench 2.1 上获得 92.8 分 (tokenstead.ai) cdnsteve 发布于 43 分钟前 | 5 分 | 隐藏 | 过往 | 收藏 | 讨论 | 帮助 准则 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime. The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.

  • Serving stack: MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers. A draft model retrained with SpecForge gives 15% longer accept lengths, and a prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20% (TTFT takes the hit).
  • Effort levels: mean steps per run 53 (medium), 80 (high), 98 (max), against 127 for SWE-1.7. Medium posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost, and lands its first real edit at a median of step 18 (SWE-1.7: 48).

Bar chart: SWE-2 medium effort needs a mean of 53 steps per run versus 127 for SWE-1.7, and reaches its first edit at a median of step 18 versus 48

Benchmarks (Cognition self-reported): FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The headline: 50.0 on FrontierCode is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3) - at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra’s. Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2’s 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.

Scorecard of Cognition's launch table: SWE-2, SWE-1.7, Kimi K3, Grok 4.6, Fable 5.1, GPT-5.6 Sol and GPT-6 Astra on FrontierCode 1.1 Main, DeepSWE 1.1, Terminal-Bench 2.1 and Terminal-Bench 4.0. SWE-2 leads Terminal-Bench 2.1 at 92.8 but trails on Terminal-Bench 4.0 at 27.3

Bar chart of FrontierCode 1.1 Main scores: GPT-6 Astra 53.3, Fable 5.1 50.9, SWE-2 50.0, Grok 4.6 48.0, GPT-5.6 Sol 47.5, Kimi K3 44.2, SWE-1.7 42.0

Proprietary weights, no local run. Cognition has not published SWE-2 weights, so there is nothing to download and no quant ladder to wait for. It is available today in Devin Desktop and CLI, with rollout on Devin Web and Fusion. Cognition publishes no per-token API for SWE-2, so the cost-per-task comparisons (64% cheaper than Fable 5.1 at FrontierCode parity) are the pricing surface, not a $/1M rate card. Every figure here is Cognition’s own number, pending independent replication.

联系我们 contact @ memedata.com