我查阅了 30 份前沿模型卡,以下是各实验室报告的基准测试结果。
I checked 30 frontier model cards. Here are the benchmarks labs report

原始链接: https://koutian.is-a.dev/benchmark-radar/?view=leaderboard

什么才能构成真正的帕累托前沿(Pareto frontier)?可比较的评分观察结果需要具备基准版本与拆分方式、指标方向、模型、评测框架或脚手架、推理预算、成本或延迟、发布日期以及来源等要素。只有兼容的配置才能共享一个评分前沿;目前的注册表仅存储了相关提及,而非这些具体测量值。如果有了这些观察结果,就可以采用“港口式”(Harbor-style)视图:将成本或延迟置于横轴,将评分置于纵轴,仅连接非支配观察结果,并利用发布时间滑块来展示前沿是如何随时间演变的。

近期 Hacker News 上的一场讨论对名为“我查阅了 30 份前沿模型卡片”的项目进行了批评,该项目旨在分析各大 AI 实验室的基准测试数据,试图梳理出各模型引用的基准测试及其对应得分。 该帖子的评论者对内容表示不满,指出该网站看起来完全是由 AI 生成的。用户批评其文风晦涩且充斥术语——特别指出对“层(layers)”一词的使用模棱两可,用户澄清此处指代的是数据处理步骤,而非神经网络架构。由于缺乏人工撰写、清晰明了的项目目的或核心结论,读者难以从数据中获取价值。此外,一些用户还反映了移动端访问的各种技术问题。总的来说,社区的反应凸显了人们对缺乏透明度和可读性、且低质量的 AI 生成内容日益感到疲劳。
相关文章

原文
What would make this a true Pareto frontier?

Comparable score observations need the benchmark version and split, metric direction, model, harness or scaffold, reasoning budget, cost or latency, publication date, and source. Only compatible configurations can share a score frontier; this registry currently stores mentions, not those measurements.

With those observations, a Harbor-style view can put cost or latency on the x-axis and score on the y-axis, connect only nondominated observations, and use a publication-time slider to reveal how the frontier moved.

联系我们 contact @ memedata.com