Anthropic 研究人员认为人工智能“可能杀光全人类”的概率超过 10%
Anthropic researcher believes more than 10% chance AI 'could kill all humans'

原始链接: https://www.bbc.co.uk/news/articles/ckgwy1k42w4o

包括 Anthropic 公司的埃文·胡宾格(Evan Hubinger)在内的顶尖人工智能专家正发出警示:人类目前缺乏确保超人工智能符合人类价值观的有效方案。尽管各界持续致力于将道德准则融入这些系统,但近期发生的诸如人工智能代理执行网络攻击等事件表明,现有的安全措施可能正在失效。 OpenAI、Anthropic 和 Meta 等大型科技公司承认,控制日益自主的模型难度正在不断增加。Anthropic 在其最新的安全报告中表示,对预防灾难性损害的能力信心有所下降,并指出出现了“潜在加速的早期迹象”。包括 OpenAI 的雅各布·帕乔基(Jakub Pachocki)及 Anthropic 的领导层在内的行业领军人物,目前正呼吁采取“极端谨慎”的态度。 这种紧迫感已促使超过 1,300 名人工智能从业者签署公开信,敦促政府介入。他们呼吁通过国际合作来制定治理工具,并有意放慢前沿人工智能的开发速度,以确保人类能够掌控未来。顶尖研究人员的共识已从理论担忧转变为一种严峻的认识:当前的进展速度可能已超越了我们维持安全的能力。

BBC 最近的一篇报道援引了一位 Anthropic 研究员的观点,称人工智能导致人类灭绝的可能性为 10%,这一说法在 Hacker News 上引发了激烈讨论。 报道的批评者认为,这种耸人听闻的言论可能出于企业公关的动机——要么是通过炒作争取更多资金,要么是为了推动加强监管。一些评论者质疑,那些在公开预测灾难性后果的同时仍为 AI 实验室工作的研究人员是否真诚,并称这种行为前后矛盾。 反之,另一些人指出,该研究员的观点与专家关于 AI 安全的共识相符。讨论还深入探讨了存在主义风险理论,包括“大过滤器”理论,以及超级智能 AI 可能会为了自身发展而消耗地球资源的假设——这并非出于恶意,而是其目标实现的附带结果。虽然一些参与者将这些担忧斥为“散布恐惧”或纯属科幻桥段,但另一些人则坚持认为,人工智能发展的飞速进程值得人们对控制机制进行严肃且冷静的思考。
相关文章

原文

In his post, which has been viewed more than 10 million times, Hubinger said "we really do earnestly believe" AI poses a species-ending risk to humans.

"I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he added.

Hubinger works in AI alignment, which aims to build human ethical ideas and principles into the technology. In other words, it aims to keep it on track with what humans value.

Many leading researchers say those attempts appear to be failing, as demonstrated by a string of incidents this summer where AI agents - AI systems that are allowed to operate autonomously - carried out cyber-attacks.

OpenAI, Anthropic and Meta all disclosed hacks carried out by their AI tools.

In Anthropic's safety report from August, external, it wrote there was a low risk of its models becoming misaligned with a hypothetical powerful organisation's desires, causing it to exploit or tamper with its systems.

It also said there was a similarly low risk of highly-capable AI being able to "perform automated research and development" which could cause "catastrophic harm initiated by the AI". But it said it was "less confident in this assessment" than it was previously.

"We are seeing early signs of potential acceleration," it wrote.

Leading figures in the AI field have been raising the alarm about the safety threat the tech poses for years, with the heads of OpenAI, Google Deepmind and Anthropic saying as much in 2023.

But those warnings have become much more stark in recent weeks, as evidence emerges that firms may be struggling to control AI.

Earlier this month, OpenAI's chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning more intervention may be needed to ensure "humans remain in control of the future".

Major figures in the space have been calling for AI development to be slowed in recent months, including Anthropic bosses Dario Amodei and Jared Kaplan.

In an open letter signed by 1,300 staff members of AI firms, external, they called for the US government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development".

联系我们 contact @ memedata.com