模型正在被蓄意变笨
Models Are Getting Dumber on Purpose

原始链接: https://w4g1.dev/blog/models-are-getting-dumber-on-purpose

人工智能实验室正日益将推理能力置于事实存储之上,这使得模型在逻辑方面表现卓越,却容易在琐碎信息上产生幻觉。这种权衡是刻意为之的:推理过程(如拆解问题)既简洁又具有永恒性,而事实知识则庞大、昂贵且极易过时。 通过将事实存储转移至外部“挂载工具”——例如网络搜索、检索系统和文档——开发者可以创造出专注于纯粹“智能程序”的模型。这种转变具有三大优势: 1. **效率**:通过剥离旨在存储事实的臃肿专家层,前沿水平的推理性能很快就能在消费级 GPU 上运行。 2. **长效**:模型不再依赖静态的训练截止日期来理解世界,因此具备了“知识保鲜”的能力。 3. **可靠性**:将答案建立在外部文档的基础上,可以将幻觉从无法修复的内部权重错误,转变为可控的数据缺陷。 归根结底,这一发展轨迹指向了未来:模型将充当灵巧的推理引擎,在运行时摄取实时、可验证的数据,从而有效地将智能与瞬息万变的事实领域分离开来。

这篇 Hacker News 的讨论探讨了“模块化 AI”的概念,即用户正从庞大的通用模型转向更小、更专业且可在本地运行的“可插拔”组件。 参与者将这种方法比作 Unix 哲学或“下载技能”,并提出了一种未来愿景:用户将特定的推理引擎与目标知识库相结合(例如,合并编程、地理信息系统和设计模块)。支持者认为,大语言模型(LLM)应减少对海量、易变事实的记忆(这常导致幻觉和数据过时),转而依赖工具调用、检索增强生成(RAG)以及基于搜索的整合。 然而,讨论也指出了几个实际障碍,包括搜索引擎日益严重的不稳定性,以及目前 AI 应用的局限性——即主要仍处于“人在回路”阶段,而非完全自主的代理工作流。最终,共识倾向于转向精简、专业的模型,使其专注于推理,同时将事实性知识交给外部的可验证来源处理。
相关文章

原文

Reasoning scores keep climbing while per-token compute keeps dropping. GLM-5.2 scores 99.2% on AIME 2026 with about 40 billion parameters active per token. Qwen3.5 scores 91.3% with 17 billion active. DeepSeek V4-Flash runs 13 billion active. For scale, GPT-4 was rumored to run around 280 billion active parameters in 2023, and it could barely solve an AIME problem. At the small end, Qwen3.5 9B fits in 6GB of VRAM quantized and roughly doubles the score of the next best model under 10B parameters on Artificial Analysis's intelligence index. If you only looked at math and code benchmarks, you'd conclude that models are getting smarter per parameter at an absurd rate.

They are, on those benchmarks. Ask the same models a plain factual question and the picture flips. On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questions. The small models barely register. Artificial Analysis measures Qwen3.5 4B and 9B at hallucination rates of 80 to 82% on its knowledge benchmark, which means that when they don't know a fact, which is most of the time, they make one up. Ask the 9B for the birth year of a minor 19th-century mathematician and you get a confident, plausible, wrong answer. The parameter count didn't drop for free. Labs are trading world knowledge for reasoning skill, and the trade is deliberate.

What the parameters were for

Facts take space. Research on knowledge capacity (the "Physics of Language Models" series has the cleanest measurements) puts it on the order of two bits of factual knowledge per parameter. If you want a model that knows the birth year of every minor Wikipedia figure, the population of every Dutch municipality, and the argument order of every function in every npm package, you pay for that in weights, and it's a big part of why frontier models grew to trillions of parameters.

Reasoning compresses much better than facts do, because it's a relatively small set of procedures applied over and over: break the problem into parts, track intermediate state, check your own work, backtrack when a step fails. Distillation and reinforcement learning on verifiable tasks turn out to transfer those procedures into small models remarkably well. Phi-4 is 14 billion parameters, trained heavily on synthetic textbook-style data, and it's good at math and bad at trivia, which tells you exactly what its training data contained. That mix used to look like a limitation of the synthetic-data approach. It now looks like the design goal.

The knowledge that survives the trade has a shape. These models are generalists: they know a little about nearly everything and almost nothing in depth. Ask one about PostgreSQL and it knows what it is, what it's good at, and roughly how MVCC works, but ask which version added a specific planner feature and you're back to invented facts. That's the right layer to keep in weights, because breadth is what lets a model understand what a question is about, know what to look up, and judge whether a source is plausible. The depth is cheap to retrieve and expensive to store, so it's the part that goes.

Facts rot, procedures don't

A frontier training run takes months and costs hundreds of millions of dollars, and the moment it finishes, the facts inside it start going stale. Library APIs change, prices change, people change jobs, and half of what a 2024 model believed about the JavaScript ecosystem was outdated before the model shipped. Every fact you bake into weights has a shelf life, and the only way to refresh it is another training run.

The procedures don't rot. Algebra worked the same way in 1970 as it does now, and so does breaking a problem down or spotting a contradiction between two sources. A model that's mostly procedure and only lightly loaded with facts doesn't age the way a knowledge-heavy model does. Its training cutoff matters much less, because the current state of the world was never supposed to live in the weights in the first place. I think this is the best argument for the whole approach: it decouples the expensive, slow artifact (the trained model) from the thing that changes daily (what's true).

The harness carries the knowledge

If the model doesn't know things, something else has to, and that something is the harness: retrieval over a knowledge base, tool calls, web search, a filesystem full of docs. I wrote earlier that Rust is a harness for agents, a source of cheap machine-checkable feedback. This is the same shape from the other side. The model contributes reasoning, and everything it reasons about gets supplied at runtime.

You can already watch agents work this way. A coding agent doesn't need to have memorized your dependency's API surface, because it greps node_modules or reads the docs before calling anything, and its answer is grounded in the version you actually have installed rather than whichever version dominated the training data. The recall that used to be a fixed cost in every forward pass became an on-demand lookup.

A frontier model on your GPU

Follow the trend a couple of years out and I think we get a model with frontier-quality reasoning, Fable-quality, that runs on a single consumer GPU. The compute half is nearly there. DeepSeek V4-Flash reasons with about 13 billion active parameters per token, well within consumer-GPU range. What doesn't fit is the other 271 billion parameters sitting in its experts, and expert layers are mostly fact storage. That's the part this whole trade makes optional. Strip the knowledge out and total size shrinks toward active size, and a 20 to 40B model at 4-bit quantization fits on the 24GB card that's been sitting in gaming PCs since 2022.

The catch is that it won't know much. Ask it a bare factual question with no tools attached and the right behavior is to say it doesn't know and go look it up. Paired with a decent harness, that's most of what I use a frontier model for today, running locally with no per-token bill and no data leaving the machine.

This mostly solves hallucination

The part I find most promising is what this does to hallucination. When a fact lives in weights, a wrong fact is unfindable and unfixable. You can't grep the weights, you can't diff them against last month, and correcting one error means a fine-tune that might break who knows what else. The model states the wrong fact with the same fluent confidence as a right one, and there's no artifact to check it against.

When the fact lives outside the model, a wrong answer has an address. The model cites a document, so you can open the document. If the document is wrong, you edit the document, and every future query gets the correction, which beats waiting for the next training run by roughly a year. Retrieval doesn't get you to zero, since a model can still misread a source or stitch two of them together wrong, but a claim with a source is checkable and a claim from weights isn't. A wrong fact in a knowledge base is an ordinary data bug, the kind we already know how to trace, fix, and write a regression test for.

There's a version of this future where the model card stops listing a knowledge cutoff at all, because what's left in the weights goes stale on a scale of years instead of weeks. The model just gets handed the world's current state at runtime, the same way a CPU gets handed a program.

联系我们 contact @ memedata.com