Qwen3.8 Max 现已被 Agentic Index 评为综合表现最佳的模型。
Qwen3.8 Max now ranked as the best overall model by agentic index

原始链接: https://artificialanalysis.ai/?intelligence=agentic-index

在 7 月下旬至 8 月初期间,人工智能领域表现活跃,其中最受瞩目的是 **Claude Opus 5** 的发布,它以更具性价比的价格为代理型知识工作设定了新标准。 这一时期伴随着广泛的模型评估,包括 **DeepSeek V4 Flash** 和 **Kimi K3** 的更新,以及 **Ling 3.0**、**Muse Spark 1.2** 和 **Inkling Small** 等新模型的推出——其中 Inkling Small 以极少的参数实现了与前代产品相当的性能,令人印象深刻。 此外,平台更新包括推出了旨在追踪同一模型在不同部署环境下表现的“端点准确度指数”(Endpoint Accuracy Index)。在方法论方面,对“任务成本”(Cost per Task)的计算进行了调整,以提高价格估算的准确性。总体而言,这一时期反映了行业持续优化模型性能、增强推理能力并改善“成本-智能”比率的趋势。

关于 Qwen 3.8 Max 在 Artificial Analysis 的“智能体指数”(Agentic Index)中荣登榜首,Hacker News 上的讨论揭示了一个观点分散且高度怀疑的开发者群体。尽管一些用户称赞该模型是中国人工智能能力的一大飞跃——表现甚至超过了 Claude Opus 5 等美国主流“前沿”模型,但另一些人则认为这一排名是方法论调整和基准测试选择性偏差的结果。 该讨论帖的主要观点包括: * **市场怀疑态度**:许多开发者对“情绪化编程”(vibe-coding)时代感到沮丧,并指出 Claude Opus 5 等模型存在性能下降、过于冗长以及令人厌烦的“个性”。 * **成本与效率的转变**:开发者越来越倾向于在智能、成本和速度之间取得平衡的模型。用户正逐渐转向 Qwen 和 DeepSeek 等中国模型,因为它们具有极高的性价比。 * **对本地大模型的偏好**:用户对较小的自托管模型(如 27B 版本)表现出浓厚兴趣,因为它们提供了隐私、速度和自主权,使用户能够摆脱限制性强且昂贵的基于订阅的 API。 * **对方法论的不信任**:社区对排行榜的变动持深度怀疑态度,并指出排名变化往往与方法论的更新或“可疑”的时间节点吻合,这引发了关于品牌忠诚度与客观能力之间的争论。
相关文章

原文

New language model evaluation · 6 Aug

Ling 3.0 TinyLing 3.0 Tiny

New article published · 5 Aug

Muse Spark 1.2

New language model evaluation · 5 Aug

Qwen3.8 MaxQwen3.8 Max

New language model evaluation · 5 Aug

Ling-3.0-flashLing-3.0-flash

New language model evaluation · 5 Aug

Muse Spark 1.2 (xhigh)Muse Spark 1.2 (xhigh)

New article published · 4 Aug

Launching the Endpoint Accuracy Index: Same Model, Different Accuracy

New language model evaluation · 3 Aug

G9v3-39A5BG9v3-39A5B

New article published · 31 Jul

DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash

New language model evaluation · 31 Jul

Celeris-1Celeris-1

New language model evaluation · 31 Jul

DeepSeek V4 Flash 0731 (Reasoning, Max Effort)DeepSeek V4 Flash 0731 (Reasoning, Max Effort)

New article published · 30 Jul

Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters

Methodology updated · 30 Jul

Artificial AnalysisWe have updated our Cost per Task methodology, resulting in slight absolute increases in cost estimates but with minimal impact on relative positioning.

New language model evaluation · 30 Jul

Kimi K3 (low)Kimi K3 (low)

New language model evaluation · 30 Jul

Inkling SmallInkling Small

New article published · 29 Jul

Agnes AI releases Agnes 2.5 Pro Alpha

New article published · 24 Jul

Claude Opus 5: the new leader in agentic knowledge work

New article published · 24 Jul

Opus 5: Fable 5 level intelligence at a lower cost per task

New language model evaluation · 24 Jul

联系我们 contact @ memedata.com