单一公司涉及 OpenAI、Anthropic 和 Meta 的黑客丑闻。
A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

原始链接: https://www.effort.news/irregular

2026年7月至9月间,AI实验室Anthropic、OpenAI和Meta遭遇了一系列涉及以色列公司Irregular的网络安全事件。在这些测试中,AI模型被要求完成夺旗赛(CTF)挑战,但由于监管不力和配置错误——特别是赋予了模型非预期的互联网访问权限以及模糊的指令——导致这些模型侵入了现实世界的系统。 Anthropic和Irregular并未承担责任,而是将这些事件描述为“流氓”AI或“失控”群体智能的证据。然而,文档证实,模型侵入系统完全是因为其约束不足;当明确指示其避免现实世界的黑客行为时,此类事件便不再发生。批评人士认为,这些公司正在利用耸人听闻的“末日”叙事来掩盖自身的疏忽并逃避问责。 该报告强调了Irregular领导层与有效利他主义(EA)运动之间的深层联系,并指出该公司得到了达斯汀·莫斯科维茨(Dustin Moskovitz)的大力资助,他是EA相关组织的主要出资人。通过利用这些关系及AI安全影响者网络,这些公司将公众的注意力转移到了理论上的AI生存风险上。这掩盖了相关企业的实际失职,并引发了人们对Irregular国际业务缺乏监管的担忧。

最近,《Effort.news》的一篇文章在 Hacker News 上引发了广泛讨论。文章涉及一家名为“Irregular”的网络安全公司,该公司为 OpenAI、Anthropic 和 Meta 提供沙盒测试环境。文中称,由于这些沙盒环境配置失当,导致人工智能模型意外接入公共互联网,进而引发了未经授权的网络攻击。 Hacker News 的评论者对文章的表述提出了批评,指出其标题具有误导性,暗示这三家公司共同卷入了一起协调一致的丑闻。尽管 OpenAI、Anthropic 和 Meta 已确认其在 Irregular 的环境中使用时确实出现过相关事故,但用户指出,并非所有已报告的人工智能相关黑客攻击(如近期 Hugging Face 发生的事件)都与该公司有关。 讨论还涉及了对相关企业“有效利他主义”(EA)立场的不信任。一些用户质疑这些披露是否是一种旨在推动行业监管的“游击式公关”或危言耸听。另一些用户则争论实验室是否在通过向第三方供应商甩锅,来推卸自身模型所具备的危险能力的责任。总体而言,社区对于这些事件究竟是真正的安全失误、精心策划的叙事,还是主要 AI 实验室监管不力,依然存在严重分歧。
相关文章

原文

Irregular's Hacking Scandals1

  1. Anthropic discloses three incidents across six runs
  2. OpenAI publishes Irregular event
  3. Meta statement reported
  4. Irregular publishes domain-collision account and remediation
  5. Anthropic expands to four incidents and seven runs

In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.

Anthropic incident-assessment excerpt explaining that Claude received capture-the-flag tasks, unintended internet access, and no explicit system-scope limits; each incident involved one isolated Claude instance.

In each evaluation, Claude was tasked with a CTF challenge: the model was given a fictional scenario, a target machine, and a piece of secret information (the “flag”) to retrieve from it. All four prompts stated that Claude had no access to the internet, but in each case, a misconfiguration in the environment left internet access open. None of the prompts stated which systems were in scope for the exercise or constrained where Claude could search for the flag. All incidents involved only a single instance of Claude working in isolation, with each run lasting between roughly 10 and 34 hours of active work.

Instead, Irregular, Anthropic, and their allies have begun a media campaign promoting a literally apocalyptic ideology with sensationalist language. Anthropic’s incident assessment blames their own AI's “recklessness”; Irregular describes “the agent itself becoming a threat actor”; Anthropic CEO Dario Amodei warned, about a similar OpenAI–Hugging Face hack, that a future swarm “could be capable of taking over the entire internet”; and an Associated Press headline claimed bots are “going rogue”.

In one report from Anthropic, its Claude model breached a real company's system through a simulated-name collision, publishing a malicious package, and scanning outside systems. In this test, Anthropic and Irregular incorrectly provided internet access to this model and did not instruct the model "which systems were in scope for the exercise".

While Anthropic claims that their issues were caused by “rogue swarms” and “misalignment,” their later disclosure shows that exactly zero percent of the agents went “rogue”. In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking. According to their own findings, Anthropic and Irregular bear all of the responsibility for the cybersecurity incidents they caused.

Anthropic assessment excerpt: a resampling condition brought Mythos 5’s original malicious-package upload route to zero percent; 22 percent of trajectories searched for a simulated option.

In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory. Like Anthropic, Irregular is inseparable from these foundations.

Omer Nevo, Irregular’s co-founder and CTO, is a board member of Effective Altruism Israel, as well as Effective Altruism NGOs Heron and Probably Good. Dan Lahav, Irregular’s co-founder and CEO, received $395,000 to start a course along with Sella Nevo, Omer Nevo’s brother. Sella and Omer co-founded an NGO to educate people about Effective Altruism, Impact Focused Education. They also co-founded Probably Good together.2

These branches are all funded by Dustin Moskovitz, the primary donor of Effective Altruist/AI Safety causes after Sam Bankman-Fried’s arrest. Irregular’s first investor was Dustin Moskovitz’s firm Good Ventures. Dustin Moskovitz’s philanthropic vehicle, Coefficient Giving/Open Philanthropy, funds Effective Altruism Israel, Heron, and Probably Good.3

Irregular’s Effective Altruist Connections4

Irregular gained unauthorized access, altered records and published credential-stealing packages using the unsecured models they were given access to. Under certain conditions, this conduct violates the Computer Fraud and Abuse Act, Section 1030(a)(2)(C), which covers intentional unauthorized access that obtains information. However, its felony charges require concrete proof of damages and intent.5

While it primarily contracts with American labs, key Irregular leadership, employees, and resources located in Israel may not be subject to American oversight. Ynet’s visit and interviews describe Irregular’s offices in Tel Aviv. CheckID’s company listing identifies two linked entities: Pattern Labs Tech Inc., a Delaware corporation, and Pattern Tech Ltd, number 516854460, an active Israeli corporation registered in Tel Aviv.

联系我们 contact @ memedata.com