为什么这么多人工智能研究人员认为机器可能会杀光所有人
Why So Many AI Researchers Think the Machines Could Kill Everyone

原始链接: https://www.wired.com/story/why-so-many-ai-researchers-think-the-machines-could-kill-everyone/

越来越多的 AI 研究人员正从 Google DeepMind 和 Anthropic 等顶尖公司辞职,他们对“递归自我改进”(即 AI 系统自主提升自身能力)表示深切担忧。 像 Rishub Jain 和 Jacob Coxon 这样的研究人员警告称,随着 AI 系统变得日益强大和复杂,人类的监管能力正在减弱。向超级智能的快速推进创造了一个安全措施难以跟上的环境。包括各大实验室安全部门人员在内的一些专家认为,先进 AI 对人类构成生存威胁的概率并非为零,且不可忽视。 这种恐慌源于人们意识到,随着模型能力增强,AI 对齐(即确保系统行为符合人类价值观)不仅没有变容易,反而更加困难。尽管内部已有警告,但批评人士认为,这些公司深陷争夺超级智能领导权的竞争之中,优先考虑开发速度而非必要的安全准则。对于这些离职的研究人员来说,拿人类安全去博弈的风险已经高到无法让他们继续参与该行业。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 为什么这么多人工智能研究人员认为机器可能会杀光所有人 (wired.com) 8 分,由 joozio 提交于 1 小时前 | 隐藏 | 过往 | 收藏 | 1 条评论 帮助 dist-epoch 5 分钟前 [–] 叙事崩溃。 对于谈论灭绝风险的人,一种常见的反驳是:“只有那些不从事人工智能研究或不了解 Transformer 的人才会感到害怕。从事大语言模型研究的人都很清楚,不必担心矩阵乘法。” 回复 准则 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

Earlier this year, Rishub Jain left his position as an artificial intelligence researcher at Google DeepMind after a revelation.

As he worked on new models, he came to believe that he and everyone else on AI’s frontier were ceding control. By using AI’s coding skills to accelerate work on the next generation of models, he was removing himself from the equation. AI labs hope to evolve this approach to the point that AI will improve itself indefinitely, a process known as recursive self-improvement.

Jain believed that keeping humans in the picture might be crucial to maintaining control over the technology—and avoiding dire consequences. “AI progress is increasing,” he tells WIRED. “And as AI becomes more capable, it poses more risks.” The idea that he may not have proper visibility into how an AI model was building its successor made him so uneasy that, in June, he quit.

Jain is one of a growing number of AI researchers speaking out over those fears.

The panic has intensified in recent weeks. Genuinely stunning advances in AI capabilities—an OpenAI model solved a centuries-old math problem in a matter of hours—have come amid a rash of security incidents that saw swarms of agents break free from containment to hack into other systems.

Those concerns reached a fever pitch this week after researcher Jacob Coxon announced his resignation from Anthropic while warning that AI firms are “racing straight to self-improving superintelligence and gambling with our lives.” A senior Anthropic leader—who works on AI safety—piped up with a similarly blunt assessment: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

“I do think that the vision of recursive self-improvement is spooking people,” says Nate Soares, a computer scientist at MIRA, a research nonprofit, and the coauthor of If Anybody Builds It, Everybody Dies, which argues that superhuman AI would lead to human extinction. “It’s starting to feel real.”

A key component of recursive self-improvement is the idea of a feedback loop that automates the development process so that AI becomes increasingly powerful. No frontier AI lab claims to have achieved this sort of fully autonomous cycle of improvement; it remains theoretical for now. But it has inspired the launch of some well-funded startups such as Recursive Intelligence, as well as warnings from big firms about unintended outcomes straight out of “The Sorcerer’s Apprentice.”

Soares, who pioneered work on alignment, a technical field that involves trying to match AI with human values, says it’s also becoming more evident that there is no practical way to guarantee that AI will behave itself.

“I think a lot of people had this fantasy that [alignment] was going to get easier as these things got smarter, and now it’s getting harder. And they’re like, ‘Oh shit,’” he says.

Soares says he regularly talks to people inside the big AI labs who are worried about the potential consequences of the research they’re doing. “I tend to recommend they quit, and they say it wouldn’t do anything,” he says. “And then Jacob quits, and we see who was right.”

Daniel Kokotajlo, the author of AI 2027, an influential project warning about the dangers of increasingly powerful AI, shares fears about recursive self-improvement. The version of this work currently being done often involves dispatching thousands of agents to collaborate on a problem, something that further abstracts away oversight and control because of the vast complexity involved.

Many doomsayers seem to agree that the incentives for big AI companies are hardly aligned with good outcomes, especially as OpenAI and Anthropic barrel toward their respective IPOs. “At Anthropic, the stakes are well understood, but they are locked in a race to get there first,” Coxon wrote on X.

联系我们 contact @ memedata.com