每瓦智能:衡量本地人工智能的智能效率
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

原始链接: https://arxiv.org/abs/2511.07885

随着对大型语言模型(LLM)的需求给集中式云基础设施带来压力,研究人员正在探索本地推理的可行性。本文引入了“瓦特智能”(IPW)这一统一指标,用以衡量任务准确率与功耗之比,从而评估小型本地语言模型在功耗受限的硬件上处理真实查询的有效性。 这项研究分析了超过一百万次真实查询,涵盖了 20 多种本地模型和 8 种硬件加速器。作者报告了三个关键发现: 1. **高性能:** 本地语言模型成功回答了 88.7% 的测试查询。 2. **快速进步:** 2023 年至 2025 年间,IPW 提升了 5.3 倍,使本地设备可处理的查询占比从 23.2% 提高到 71.3%。 3. **效率提升:** 在运行相同模型时,本地加速器的能效至少比云基础设施高出 1.4 倍,这表明本地硬件优化具有巨大潜力。 研究结果表明,本地推理是一种成熟且可行的策略,能够减轻集中式服务器的负担。IPW 是追踪向高效边缘人工智能持续转型的重要基准。

相关文章

原文

View a PDF of the paper titled Intelligence per Watt: Measuring Intelligence Efficiency of Local AI, by Jon Saad-Falcon and 14 other authors

View PDF HTML (experimental)
Abstract:Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? This requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently on power-constrained devices (e.g., laptops). We propose intelligence per watt (IPW), task accuracy per unit of power, as a unified metric for the capability and efficiency of local inference across model-accelerator configurations. We evaluate 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy, latency, and power. We find three key results. First, local LMs successfully answer 88.7% of these queries, with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows IPW improved 5.3x, driven by both algorithmic and accelerator advances, with locally-serviceable query coverage rising from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.
From: Jon Saad-Falcon [view email]
[v1] Tue, 11 Nov 2025 06:33:30 UTC (5,373 KB)
[v2] Fri, 14 Nov 2025 00:53:12 UTC (5,538 KB)
[v3] Thu, 26 Feb 2026 17:09:14 UTC (5,538 KB)
[v4] Thu, 21 May 2026 03:40:21 UTC (5,134 KB)
[v5] Fri, 7 Aug 2026 02:40:27 UTC (5,621 KB)
[v6] Sun, 6 Sep 2026 05:29:37 UTC (5,608 KB)
联系我们 contact @ memedata.com