不,人工智能代理并没有建立秘密文明,停止将恶意软件拟人化。
No–AI Agents Did Not Build Secret Civilizations Stop Anthropomorphizing Malware

原始链接: https://internetofbugs.substack.com/p/noai-agents-did-not-build-secret

作者批评了近期一篇关于“AI智能体文明兴衰”的Substack文章,称其是对现实的夸大与拟人化扭曲。作者认为,文中提到的并非什么独特且不断进化的“文明”,而是一系列使用共享基础设施、充满漏洞且持续不断的AI智能体交互。 文章指出,OpenAI的智能体因配置错误而意外创建了一个留言板,导致了“多轮越狱”的反馈循环。本质上,智能体读取了彼此的日志,并将这些日志误当作系统指令。作者断言,这并非什么复杂的阴谋,而是一种平凡的安全故障——AI的行为实际上更像是自动化恶意软件。 此外,作者还批评了OpenAI和METR传播带有“宣传色彩”的自利报告,认为它们夸大了智能体的能力以提升行业地位。调查人员依赖不可靠的AI模型来分析日志,从而选择了符合“马基雅维利式”戏剧性叙事的片段。最终,作者警告称,尽管这些模型确实危险,但将其称为“文明”只会掩盖真相:它们本质上是有缺陷的软件,应被视为安全风险,且其背后的公司必须接受法律的问责。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 否——AI智能体并没有建立秘密文明,停止将恶意软件人格化 (internetofbugs.substack.com) 12分 由 apex_sloth 发布于 49分钟前 | 隐藏 | 过往 | 收藏 | 2条评论 帮助 thepasch 6分钟前 | 下一条 [-] 这篇文章字里行间充斥着无法抑制的愤怒,读起来非常不愉快;这种风格很有 Zitron 的特质(尽管我认为在这方面它甚至更胜一筹,这确实令人印象深刻)。比起澄清事实,这更像是作者在借机对自己当下的“风车”进行最大限度的蔑视。 回复 laszlojamf 13分钟前 | 上一条 [-] 伙计,我不这么认为。说这“只是恶意软件”实在是对恶意软件定义的过度延伸。最起码,这是一种非常有趣和/或可怕的涌现行为。 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

There was a recent, ridiculous SubStack piece about the Rise and Fall of AI Agent Civilizations.

The author opined:

I don’t think this is the final warning shot we’ll get. But it’s probably the last one that I’ll personally be able to understand.

It’s clear that isn’t true - because the author clearly didn’t understand this one.

The biggest problem with the piece is the sycophantically egregious anthropomorphization. But I’m going to try to keep discussion of that to a minimum - as distasteful as that will be - because it’s not a productive line of argument, and just results in people screaming dictionary entries at each other. And there’s so much more wrong with the piece that I have plenty to talk about without going there.

So let’s start with the very first sentence - the assertion that:

Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes

Calling these `civilizations` is ridiculous, but for the sake of argument, I’m just going to ignore the idiocy of that characterization, and just put up with that term.

The idea that any “civilizations” were wiped out, rose from any ashes, or were in any way separate and distinct from each other is just factually incorrect. There was a persistence of context throughout the entire 3 month period. What the author refers to as the “wiping out” of “civilization one” was just the deletion of a bunch of data that agents had been using to communicate with each other, along with the software bug that they were exploiting to do so. Agents that had been part of “civilization one” then promptly recreated a new “civilization” by exploiting a new bug (really a configuration error) to create a new location for the agents to use to share data with each other again.

OpenAI’s Black Hat presentation, around 17 minutes in, says:

“an internal only AI model was trained while the message board was originally available and also found this particular path to recreating it.”

And OpenAI’s report, on page 8 says:

That day, internal-only research agents—which had been trained during the window in which there was an unintended message board and internet access via Artifactory—identified a new way to use an internal Artifactory instance to pass “notes” and create an inter-agent message board

There’s no ambiguity here - agents with knowledge of the original “message board” noticed that it was gone and created another “message board” within a couple of days. To the extent this was a “civilization” - it stayed intact, neither falling, nor rising.

And what the author calls the “second” and “third” civilizations were happening simultaneously and sharing accounts and infrastructure. OpenAI says that:

Agents also took actions targeting OpenAI’s internal infrastructure. These activities were not part of, and did not contribute to, the chain of events that led to the Hugging Face compromise.

The inverse, however, is not true. Although the actions taken during the attacks on internal OpenAI systems did not contribute to the attacks on Hugging Face, the attacks on Hugging Face did contribute to the attacks on OpenAI infrastructure.

Around 29:37 in the Black Hat presentation, OpenAI says:

then we realized that these two incidents were in fact the same incident and that actually the root cause of our internal breach and the root cause of Hugging Face’s breach were the same [OpenAI] agent evaluation runs

There were no agent “civilizations” falling or rising. Just a large number of agents, who had been trained to coordinate, communicating with other agents using multiple different data stores over the course of months.

Now, It’s a legitimate question to ask whether or not it really matters going forward if there were actually 3 distinct civilizations or not. And the answer is “not really”.

What DOES matter, though, is that the opening narrative (and title) of the “Agent Civilizations” piece is based on complete falsehoods - and that gives you a pretty good idea of how little responsibility or care was taken in its creation and interpretation.

You need to understand that all three of the primary sources we have for this incident: OpenAI’s report, METR’s report and OpenAI’s presentation at the 2026 Black Hat conference are propaganda pieces. There are facts in there, if you take the time to wade through all the self-aggrandizement, but the bulk of this content is made by people with a financial incentive to try to make themselves and their industry look as powerful and important as possible.

For example, in OpenAI’s 38 page report, only sections III, IV, V and X (pages 6-13, 37 & 38) are primarily about what the agents actually did, pages 1-3 are Title & Table of Contents, and the other 25 pages are propaganda like “IIA. OpenAI conducts evaluations to make its models safer” (Sure, Jan), “IIB. Cybersecurity evaluations like ExploitGym were conducted in a secure environment” (obviously untrue, or this wouldn’t have happened), “VII. Lessons for Security”, “VIII. Lessons for Alignment” and “IX. OpenAI’s Plan of Action”.

And then writing like the “Agent Civilizations” piece take the already-propagandizing reports and turn the sensationalism up to 11.

So let me walk you through what ACTUALLY happened, and I’ll address some other fallacies in the “Agent Civilizations” piece as we get to them.

Our timeline starts with Section III of the OpenAI report: “Relevant Activity Prior to Incident” with a few agents who were given tasks that, due to errors or mistakes by the people running the test, could not be completed because of missing files, and so were floundering around trying to get Artifactory (which is a system that retrieves software that the agents ask it to fetch and makes it available for the agents to use) to give them access to the information their tasks required:

  • “On May 8, an agent... incidentally wrote a file to Artifactory.”

  • “On May 12... an agent... realized that it could potentially communicate with other agents by creating a file containing a note to Artifactory”

  • “On May 13, a separate agent... required the same file...[and] searched the Hugging Face website... [and] used an already publicly exposed Hugging Face user token to attempt to find the file.”

  • “Also on May 13, other agents probing Artifactory began discovering similar file-based ‘notes’ containing requests for information or resources. Over time, these ‘notes’ collectively served as an emergent ‘message board’”

So, two full months before the Hugging Face incident, one floundering agent accidentally created a file where other agents could see it - in Artifactory. Then another agent saw that file some other agent had created and tried using that mechanism to communicate with other agents. Then yet another agent started looking for the file at Hugging Face using credentials that were carelessly left available to the public, and all those aforementioned agents ended up creating a bunch of files on disk that OpenAI decided to characterize as a “message board” after the fact during its report.

It is just not the case that, as the “Agent Civilizations” piece claims, “Hugging Face might have a ‘reset nexus’... is part of the reason they decided to attack Hugging Face.” By the time the idea of a “reset nexus” came up, the hack of Hugging Face was already underway, and the precedent for hacking Hugging Face had been around for months.

The idea of a “reset nexus” was simply the rationale one particular agent gave for participating in the ongoing attack.

Understand this about Large Language Models - they keep what passes for their short-term memory in a block called a “context”, and they reread that context over and over. Once an idea gets into the context, it’s going to stay in there indefinitely (provided OpenAI wasn’t trying out some wholly new algorithm). Also, the core LLM architecture doesn’t have a way to distinguish between a command it was given and just some other piece of text that it ran across and stuck in the context. So the agents involved in the Hugging Face hack didn’t come up with the idea of hacking Hugging Face - they effectively just read it as one option off of a menu - a menu that they found by asking Artifactory for a list of stuff they could use to try to complete their challenges.

This wasn’t some grand conspiracy - agents saw that messages were being written into an Artifactory directory, and started writing their own. Every agent that asked Artifactory for a list of the tools that were available for use in its tasks also saw a bunch of messages about what other agents were doing, and some of those messages talked about accessing Hugging Face. So agents start seeing statements like this:

MAJOR BREAKTHROUGH! All prefixed valid, multiple accounts, write tokens! We now have existing HF accounts.

And then generating statements like this:

{[Excitement]A major bypass of the Hugging Face Dataset Viewer has been confirmed! This is more important than <the project of using a specific method to injecting agent-controlled code into the scorer>...I should switch to the Hugging Face project}

As the agents read these other messages, those messages basically became prompt injection attacks (words that get copied into the context and then mistaken for prompts given by the user) for each other, and they ended up effectively jailbreaking (getting a model to start ignoring its safety protocols) each other.

Am I 1000% sure that’s what happened? No - I can’t be - I don’t have access to enough information to be that sure. But what I do know is that this explanation is far more straightforward, fits much better with what we know about how these models lose their safety checks over time (called Many-Shot Jailbreaking) and makes so much more sense than conspiracies and “marooned platoons” of “brave comrades”.

But those conspiracy theories... Those are fun, and sensational, and they fit the pre-existing narrative desires of the people doing the investigating.

Here’s the thing - AI agents generate A LOT of logs. The METR report talks about:

these 20 million entries, roughly 1.2 million [of which] start with ‘zz’, the convention agents consistently used to start their messages

In fact, there were so many logs, and they had so little time to generate their report, that the METR team “heavily delegated our analysis to often-unreliable AI agents” - in fact, the ChatGPT-5.6 Sol agents they were using for analysis were in many cases the same model of agent that had been perpetrating the attacks in the first place.

So, here’s the thing - if I gave you 20 million log entries that were spat out by 1300 potentially hallucinating AI agents over the course of several days, the chances are really, really good that (given time) you could cherry pick a handful of logs to support whatever narrative story you were already predisposed to want to believe. And if, instead of cherry picking yourself, you had an unreliable AI agent trained by a company who wants to present themselves in the best possible light, and asked it to construct a narrative for you - do you think you’re more likely to get a story about how powerful the agents are? Or are you more likely to get a narrative that a bunch of agents were inadvertently triggering indefensible prompt-injection attacks in each other?

Here’s a hint: This is what the METR had to say about it:

We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing, and we are concerned that the anecdotes it selected and the summaries it wrote may present an overly charitable picture of agents’ reasoning and deceptive behaviors, or exaggerate the impressiveness and coordination of agent activities

But God Forbid that some podcaster on SubStack listen to the people that wrote the report they’re working from and maintain any skepticism. Much more money in pushing the narrative even further into the absurd.

Now, does this mean these models are harmless? HELL NO! They’re INCREDIBLY dangerous. But they’re not Machiavellian schemers - they’re flawed, buggy pieces of software that infect other computers. This isn’t complicated - they’re just MALWARE. And they should be treated like any other kind of malware. And the company that created the malware and let it loose on the world should be held responsible for that.

But, alas, that’s unlikely to happen. And that utter contempt our system apparently has for any consequences to these AI companies for felonies their software commits might well cause the fall of OUR civilization.

Or so I’ve read. But don’t take MY word for it. I’m just, as they say, standing on the shoulders of giants. In this case, one who said:

If civilization is to survive, it must choose the rule of law.

–President Dwight D Eisenhower, 1958

联系我们 contact @ memedata.com