OpenAI 的新推理技术引发了 AI 安全专家的担忧
OpenAI's new reasoning technique alarms AI safety experts

原始链接: https://techcrunch.com/2026/09/02/openais-new-reasoning-technique-alarms-ai-safety-experts/

据报道,OpenAI 的新型“Astra”模型采用了一种被称为“递归深度”或“不透明递归”的技术,该技术允许 AI 通过非线性的循环迭代来处理查询。与提供顺序且清晰可读的“思维链”(CoT)的标准推理模型不同,这种方法留下的可追踪步骤较少,这引起了 AI 安全专家们的极大担忧。 包括来自 Redwood 的研究人员在内的批评者和行业倡导者认为,这种转变威胁到了监控 AI 是否存在失控或危险行为的能力。由于不透明递归主要发生在模型的潜空间内,它可能会掩盖推理过程,导致难以审计 AI 得出结论的方式。专家担心,如果这种做法成为行业标准,可能会引发一场以牺牲透明度来换取性能的“逐底竞争”。 尽管 OpenAI 坚称 Astra 目前对该技术的使用受到限制,且仍致力于可读的思维链监控,但该报告已引发了整个行业的广泛焦虑。随着谷歌 DeepMind 和 Anthropic 等竞争对手据传也在讨论类似的方法,安全专家警告称,扩展此类架构最终可能导致 AI 的决策过程变得完全不透明且无法监控。

相关文章

原文

OpenAI’s new Astra model will use a reasoning technique called “recurrent depth” that allows it to operate outside of the sequential thinking that characterizes most reasoning models, The Information reported on Tuesday. This technique, also called “opaque recurrence,” will likely make the model’s chain of thought more difficult to monitor — and that has AI safety experts rattled.

While Astra’s use of the technique is reportedly limited, its emergence has still raised significant concerns among AI safety experts.

“I am extremely concerned by the reporting that Astra uses opaque recurrence,” wrote Redwood CEO Buck Shlegeris in a post after the news broke. “I don’t know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroys CoT monitorability.”

Longtime AI safety advocate Zvi Mowshowitz also weighed in and wrote that laws might be necessary to prevent a “race to the bottom” among AI labs. 

“The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can,” Mowshowitz wrote. “More intensive use of such techniques would probably damage monitorability.”

Under normal circumstances, a reasoning model’s chain of thought provides the sequential steps taken by the model as it attempts to solve a problem. While the representation is imperfect, it still serves as a valuable tool for monitoring misbehavior or misalignment. In the case of OpenAI’s recent rogue agent activity, chain-of-thought records were an important tool in teasing out why agents behaved the way they did.

In opaque recurrence, the model takes a less linear approach, processing the same query several times in a loop. The result leaves fewer legible traces, effectively side-stepping a conventional chain-of-thought record.

Crucially, Astra’s use of the technique appears to be limited. The model’s chain of thought is still expected to be legible, and the company pushed back against any suggestion that it would shift to “neuralese.” OpenAI has already announced plans for extensive chain-of-thought monitoring systems as part of its forward-looking safety plans.

In a post on X, OpenAI chief scientist Jakub Pachocki emphasized the lab’s commitment to legible chains of thought. “OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote. “It’s a core goal of our current research program.

All AI models do some quantity of opaque reasoning, and few researchers take chain-of-thought logs as a direct representation of a model’s reasoning. Still, those caveats don’t dispel the concern that opaque recurrence may make AI reasoning harder to monitor, particularly as it grows in use across different models. In a follow-up report Wednesday morning, The Information reported that both Anthropic and Google DeepMind were already discussing the technique.

In a post responding to the news, Redwood Research chief scientist Ryan Greenblatt said opaque reasoning could easily scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels.

“My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space,” Greenblatt wrote. “I hope it isn’t too late to avoid the most concerning architectures and that OpenAI will stop here.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

联系我们 contact @ memedata.com