Gemini 4 阿贡
Gemini 4 Argon

原始链接: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

Gemini 4 Argon 被定位为一款高能力模型,面向需要长期执行的多步骤企业任务,兼具编程、推理和多模态理解能力。据 Google 工程师表示,他们使用该模型进行调试、大规模代码库迁移和算法设计;其在 DeepSWE v1.1 上的成绩据称达到 77.9%,处于领先水平。它还位居 GDP 加权 Vals 指数榜首,并在金融和法律智能体基准测试中表现突出。 Argon 以 51.3% 的成绩位列 Zapier AutomationBench 第一,展示了强大的端到端业务自动化能力。其多模态能力包括专业图表分析、长视频理解以及处理系列文档 reportedly 在 LVBench 上取得了 91.7% 的领先成绩。 总体而言,Argon 强调长期任务执行、跨领域知识工作、视觉推理和实用的企业自动化。

讨论聚焦于 Google 的 Gemini 4 Argon。这是一款前沿模型,最初仅向部分网络安全防御人员开放,而非普通订阅用户。Google 声称,该模型在基准测试中表现强劲,并取得多项内部成果:能够让智能体将大型 C/C++ 代码库迁移至 Rust,其中包括 Fuchsia 中拥有 80 万行代码的 Zircon 内核;此外,还能优化数据中心内存,并改进量子电路。首期 API 的输入和输出价格分别为每百万 token 2 美元和 10 美元,之后将上涨至 4 美元和 20 美元。 评论者意见不一。支持者认为,Argon 表明 AI 正在实现跨越式发展,加快编码、科研和逆向工程的发展速度,同时削弱科技公司的“护城河”。批评者则质疑基准测试是否已经趋于饱和、实际使用成本是否过高、Google 延迟开放该模型的原因、割裂的产品界面,以及 Antigravity 薄弱的权限管理机制。 Gemini 3.8 Flash 在编程和系统管理方面受到好评,但用户也报告了幻觉、上下文丢失、危险的数据库操作以及对齐失败等问题。更广泛的讨论集中在以下几个方面:AI 是否会成为由少数超大规模云服务商主导的商品;开源模型能否维持竞争;以及由智能体驱动的 Rust 重写能否安全地取代成熟的 C++ 系统。
相关文章

原文

Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.

Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.

Beyond coding, Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.

Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.

联系我们 contact @ memedata.com