双子座 3.7 闪电版
Gemini 3.7 Flash

原始链接: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/

Gemini 3.7 Flash 在以下三个关键领域较 3.6 Flash 展现出了显著的性能提升: * **编程与开发:** 该模型在调试、问题解决以及生产级代码生成方面取得了长足进步,在 FrontierCode(43.6% 对比 34.4%)和 DeepSWE(65.3% 对比 49.0%)基准测试中表现尤为突出。 * **Web 开发:** 在 UI 生成方面表现卓越,具有更佳的设计还原度和功能性布局创建能力。其优势体现在更高的 WebDev Arena Elo 分数(1588 对比 1538)上。 * **复杂推理:** 在金融、法律和生物科学等专业领域,3.7 Flash 展示了更强的推理能力和准确性。在文档处理任务中,它大幅超越了前代产品(GDP.pdf 测试中为 34.0% 对比 22.0%),并在通过 AutomationBench 执行实际业务工作流时表现出更高的效率(30.4% 对比 17.0%)。 总而言之,Gemini 3.7 Flash 在准确性、技术能力和实用工作流自动化方面均实现了显著进步。

关于 Gemini 3.7 Flash 发布的 Hacker News 讨论呈现出褒贬不一的局面。虽然一些用户认可该模型基准测试的提升——指出其相比 3.6 Flash 具有更高的“人工分析”得分和性价比——但也有人质疑其价值主张。 批评者认为,对于纯文本任务,DeepSeek V4 和 Luna 等竞争模型能以相当的性能提供显著更低的成本。一些参与者还强调,近期发布的如 Grok 4.6 等模型在性能上似乎优于 Gemini 3.7 Flash,且价格更具优势,这引发了人们对谷歌是否仍处于行业前沿的怀疑。 相反,一些用户表示,Gemini 3.7 Flash 在特定工作负载下仍保持着优于竞争对手的速度优势。总体而言,社区正在讨论谷歌是在“前沿”领域有效竞争,还是正在退化为一个主要专注于快速、中等能力模型的细分领域,一些评论者对未能发布新的“Pro”级产品表示失望。
Hacker News 最新 | 往日 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 [重复] Gemini 3.7 Flash (blog.google) 46 积分,由 meetpateltech 于 1 小时前发布 | 隐藏 | 往日 | 收藏 | 1 条评论 帮助 tomhow 4 分钟前 [–] 评论已移至 https://news.ycombinator.com/item?id=49289112。 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution. It also achieves higher first-pass code accuracy and has improved performance in generating production-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).

In web development, 3.7 Flash generates more functional layouts and feature-complete apps in fewer prompts. For UI generation, the model shows high design adherence and parity based on a reference input, whether it’s a screenshot, an image, or a full design system. It outperforms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowledge-dense fields like finance, law, and biosciences, 3.7 Flash delivers improved reasoning and accuracy. It significantly outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%), an eval for testing a model’s ability to process complex documents. It also surpasses 3.6 Flash in AutomationBench, demonstrating it can more effectively complete real-world business workflows (30.4% vs 17.0%).

联系我们 contact @ memedata.com