```人工智能分析指数 v4.2```
Artificial Analysis Intelligence Index v4.2

原始链接: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2

Artificial Analysis 发布了 Intelligence Index v4.2 版本。这是一个过渡性更新,旨在跟上人工智能快速发展的步伐,为 v5 版本的正式发布做准备。该版本引入了更严苛、更真实的测试任务,并将私有预留测试集的权重提高至 40%,以防止模型出现“刷分”现象。 主要新增内容包括: * **AA-Briefcase**:一项针对代理型知识工作的评估,涉及复杂的、为期数周的项目。 * **GDP.pdf**:一项专业文档推理测试,要求模型整合 4,592 页图表、表格和文本中的数据。 * **评分增强**:升级了基础设施,包括改进采样方式并重新锚定 Elo 分级,以确保更高的稳健性和评分准确性。 目前的排名显示,Anthropic 的 Claude Fable 5.1 在总指数中领先,OpenAI 的 GPT-6 Astra 紧随其后,并表现出显著的效率提升。Anthropic、OpenAI、Meta 和 Z.AI 目前代表了“单次任务成本”的前沿水平。此次更新强化了该指数对实际应用价值的关注,随着团队继续推进 v5 版本的全面开发,未来还将发布更多增量更新。

Hacker News 最新 | 往日 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 Artificial Analysis Intelligence Index v4.2 (artificialanalysis.ai) 6 分,由 nojs 于 27 分钟前发布 | 隐藏 | 往日 | 收藏 | 讨论 | 帮助 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming

Intelligence Index v4.2 changelog:

+ AA-Briefcase, our agentic knowledge work evaluation with a private test set

+ Surge’s GDP.pdf, long context document reasoning across 4,592 PDF pages

- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated

… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness

This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January.

We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.

Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!

Intelligence Index v4.2 changes in detail:

➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

➤ Adding GDP.pdf: Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.

➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.

➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.

Key results:

➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z.AI occupy the updated Cost per Task frontier.

➤ GPT-6 Astra dominates the output token frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)

GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)

Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points

AA-Briefcase is our frontier in-house evaluation with a private held-out test set. The evaluation tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

OpenAI leads GDP.pdf with GPT-6 Astra at 33.2% and GPT-5.6 Sol at 28.2%, followed by Claude Fable 5.1 at 26.2%

Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied

Full per-model breakdowns below:

Read more about Artificial Analysis Intelligence Index v4.2 at https://artificialanalysis.ai/methodology/intelligence-benchmarking

联系我们 contact @ memedata.com