SpaceXAI 的 Grok 4.6 在人工智能分析指数中得分为 61
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

原始链接: https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis

Grok 4.6 已达到人工智能分析指数的前沿水平,得分为 61 分,仅次于 Claude Opus 5,并与 GPT-5.6 Sol 持平。相较于前代产品 Grok 4.5,其得分提升了 5 分。 该模型最突出的优势在于其代理(Agentic)性能。它在现实任务中表现优异,在终端操作和复杂的、多轮客户服务模拟方面跻身顶级模型之列。在长周期知识型工作中,Grok 4.6 达到了 Fable 5 级 Elo(1577 分),同时展现了卓越的效率;其完成任务所需的轮次约为 Claude Opus 5 的一半,输入标记(Tokens)用量仅为其四分之一。 至关重要的是,Grok 4.6 保持了其具有竞争力的定价(每百万标记 2 美元/6 美元),使其处于智能与成本的帕累托前沿。通过以竞争对手几分之一的成本提供前沿级别的推理和代理能力,它为繁重的推理工作负载提供了巨大的价值。凭借 50 万的上下文窗口和改进的架构效率,Grok 4.6 巩固了 SpaceXAI 作为人工智能领域顶级竞争者的地位。

抱歉。
相关文章

原文

Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic.

Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3

➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models

➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier

➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)

Other model details: ➤ Context window of 500k tokens (unchanged from Grok 4.5)

➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits

Agentic performance

Grok 4.6's strongest results are on agentic work rather than static reasoning. On GDPval-AA v2, our leading measure of real-world agentic knowledge work, it scores an Elo of 1753 - behind only Claude Opus 5, and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max given overlapping confidence intervals.

The pattern holds across task types. 𝜏³-Banking (50.7%) tests multi-turn customer service with tool use and places Grok 4.6 in the top two, while Terminal-Bench v2.1 (88.4%) puts it level with the leaders on terminal-based software tasks. Few models are simultaneously competitive across knowledge work, customer service and terminal use; combined with its pricing, this places Grok 4.6 on the cost vs. performance Pareto frontier for every agentic evaluation in the Intelligence Index.

Cost

Holding headline pricing flat across a generation is unusual at the frontier, where intelligence gains have typically been accompanied by price increases. Grok 4.6 delivers a 5-point Intelligence Index gain at unchanged $2/$6 pricing, and our measured cost per task of $0.84 reflects both that pricing and reasonable token efficiency.

The comparison that matters for buyers is against the models scoring within two points of it: Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. Grok 4.6 offers effectively the same Intelligence Index score as GPT-5.6 Sol at a fraction of the output token price, which is the dimension that dominates cost in reasoning-heavy workloads.

Long-horizon knowledge work

Grok 4.6 debuts on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577. This places it at Fable 5-tier, behind the Claude Opus 5 family, with consistently strong performance across rubric grading, presentation quality and analytical quality rather than strength in one dimension offsetting weakness in another.

The efficiency profile is as notable as the score. Grok 4.6 resolves tasks in ~53 turns and ~0.5B input tokens on average, against ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). Long-horizon agentic work accumulates context rapidly, so a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage well beyond its per-token pricing.

Full results

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks of Grok 4.6: https://artificialanalysis.ai/models/grok-4-6

联系我们 contact @ memedata.com