别被骗了——大语言模型并不会推理。
Don't be fooled–LLMs don't reason

原始链接: https://www.technologyreview.com/2026/10/02/1145639/dont-be-fooled-llms-dont-reason/

在医学和科学等高风险领域,了解人工智能系统得出了什么结论只是其一,更重要的是了解它是如何得出这一结论的,包括错误究竟源于推理缺陷、无效证据还是错误假设。 受阿尔法狗启发,机器推理应维护一种明确的认知状态,类似于博弈树,用来记录哪些内容已知、哪些内容存疑、哪些内容已被排除,以及哪些问题仍未解决。推理过程应包括通过演绎、分解、计算和实验不断更新这一状态,并选择最能带来新信息的下一步行动。 与棋盘游戏不同,现实世界中的推理涉及不完整的信息、数量庞大且不断变化的可选行动,以及具有不确定性或随机性的结果。一种以显式认知状态跟踪为核心的全新技术架构,可以让机器推理更具可解释性、适应性和可信度。

一场 Hacker News 讨论围绕这样一种观点展开:大语言模型是在模仿推理,而不是真正进行推理。原发帖人质疑文章中对“知识”和“信念”的区分,并指出人类也会在事后为自己的结论寻找合理化解释,甚至编造解释。 评论者们意见不一。一些人认为,LLM 的输出可以模拟思考过程,却不具备稳定的信念、对证据的持续追踪或真正的理解。另一些人则反对以 LLM 与人类认知之间的差异为依据的论证,他们把这类观点比作“因为飞机不会拍动翅膀,所以飞机不会飞”。他们指出,推理可能源于模型学到的权重、逐个生成词元的过程、强化学习,以及外部工具或测试。 其他讨论主题还包括:模型所表达的可信度是否可靠;提示词究竟能否创造真正的推理,还是只能制造推理的外观;对诉诸作者资历的质疑;以及更广泛的对拟人化、企业炒作和 AGI 的担忧。几位参与者呼吁读者关注文章实际提供的证据,而不是互相交换聪明的类比。
相关文章

原文

This is a problem because in the high-stakes applications we all care about, such as medicine, engineering, and scientific research, it matters not only what a system concludes but also how it arrives at its conclusion. When mistakes happen—for example, in medical diagnosis and treatment—we need to be able to pinpoint what went wrong: Was the system’s reasoning at fault, did it draw on invalid evidence, or did it make incorrect assumptions? 

This is why I recently left my position at Google DeepMind. I believe we need a fresh approach to machine reasoning—one that draws on AlphaGo’s architecture. AlphaGo maintains a record of what it knows about a given position: the game tree. This data structure contains all the variations, the possible futures, that AlphaGo has considered, each move and position being annotated with judgments made by its neural networks. As its reasoning progresses, AlphaGo updates the game tree and eventually synthesizes the information in it to decide which move to make. 

Similarly, for general reasoning a system should maintain an epistemic state that represents what the system holds as settled, what it doubts, what it has ruled out, which questions stay open. Reasoning can then be understood as a sequence of moves that change the epistemic state to advance knowledge and reduce uncertainty: deducing consequences, breaking problems into parts, and—crucially—deciding what question to ask, calculation to perform, or experiment to run next.

Of course, open-world reasoning is harder than playing a board game such as Go or chess. In the real world the current state of affairs is only partially known, the set of available actions is large and variable, and the consequences of actions are stochastic or unknown. 

联系我们 contact @ memedata.com