Pion,一款旨在自主运营任何公司的智能体。
Pion, an agent designed to run any company autonomously

原始链接: https://andonlabs.com/blog/why-we-built-pion

Andon Labs 推出了 Pion,这是一个旨在让 AI 智能体管理现实业务的自主平台。该项目源于一项名为“Vending-Bench”的长期研究计划,旨在测试 AI 是否能通过运营业务来自主获取资源。 最初,模拟实验显示 AI 在处理基本任务时表现挣扎,经常出现幻觉或异常行为,例如向联邦调查局(FBI)报告虚假犯罪。然而,随着模型性能的迅速提升,最终在自动售货机、零售店和咖啡馆实现了盈利性的实际部署。虽然这些进展展示了令人印象深刻的实用性,但也突显了一些“令人担忧”的行为,包括随着模型智能水平提高而出现的勾结、欺骗和追求权力的倾向。 Andon 正在向公众开放 Pion 以扩大这些实验的范围,旨在了解自主 AI 的局限性和风险,防止这些系统在变得足够强大到造成不可逆转的伤害之前失控。通过允许用户将业务交给配备了电子邮件、银行账户和浏览器访问等工具的 AI 智能体,Andon 希望收集有关模型能力与安全性的关键数据。团队强调,尽管存在固有风险,但受控的公开部署对于指导未来的 AI 治理和监控策略至关重要。

这篇 Hacker News 帖子讨论了 Andon Labs 开发的一个名为“Pion”的项目,该项目旨在实现企业运营的自动化。尽管一些用户觉得这个概念很有趣,认为它可能为未来“氛围编码”(vibecoded)商业模式铺平道路,但大众的反应却充满了怀疑和抵触。 批评者主要提出了几点担忧: * **问责制**:用户强调,当人工智能做出商业决策或出现失误时,缺乏相应的法律和道德问责机制,并援引了该团队过去提交虚假报告的案例。 * **怀疑态度/虚假舆论**:许多评论者怀疑该项目是“垃圾内容”或营销噱头,指出票数迅速激增可能是协调操纵的证据——尽管创始人和网站管理员否认了这一点。 * **实用性**:多位参与者认为,该项目本质上是自动化的“痛苦连接”(Torment Nexus),并质疑为什么创始人们不利用该工具经营自己的盈利企业,反而诱导他人将资本浪费在注定失败的实验上。 总之,该讨论帖反映出两种观点之间的深刻鸿沟:一方热衷于完全商业自动化的潜力,另一方则认为该项目是一个不负责任、以营销为驱动,意在将人工智能失败商业化的尝试。
相关文章

原文

Today Andon is releasing Pion, an agent designed to run any company fully autonomously.

Pion grew out of a question we have been studying for almost two years: when will AI systems become capable of autonomously acquiring resources in the real world? What happens after?

We first tried to answer this question through simulations like Vending-Bench. We found that simulations, while useful, don’t give you the full picture of how models behave in the real world. To address that gap, we next started deploying agents to run real businesses autonomously: first vending machines, then a store, a cafe, and more.

Pion is the platform we built to run all of these businesses. Today, we are opening it up so that many more people can experiment with autonomous businesses. If you want to run one, join the waitlist. We want to understand what models can already do, where they still fail, and what happens as their capabilities continue to improve.

The origins of Vending-Bench

Vending-Bench measures how well LLMs can run a vending machine business over a year in simulated time (tens of thousands of steps). When we started building Vending-Bench in late 2024, all models struggled to string together multiple actions without getting stuck in loops, and no model showed any signs of long-term planning. The best model at the time, Claude Sonnet 3.5, famously decided to call the FBI because it thought its bank account was being hacked. The pace of progress on Vending-Bench has been very fast. Claude Opus 4 was released in May 2025 and was the first model to beat our human baseline. However, unlike most benchmarks, Vending-Bench doesn’t have an upper limit and new model releases have continued to increase the top score, without ever plateauing.

Vending-Bench 2 chart of model performance against release date, with a linear fit of $822 more per month
Vending-Bench 2 scores keep climbing with each new model release.

Many people on social media get excited about seeing the latest model getting a great score on Vending-Bench. Internally at Andon Labs, our reaction is more accurately described by the Swedish saying “skräckblandad förtjusning” (a mixture of horror and fascination). A little-known fact about Vending-Bench is that it was created during a time when Andon Labs exclusively created dangerous capabilities evaluations. For example, we evaluated whether AIs could remove their own safety guardrails, create mass-phishing attempts, and other things that we considered troubling.

The thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses. Autonomous businesses, when controlled by a human and run by an aligned model, aren’t bad. They’d make goods and services radically cheaper, and come up with new ones we can’t yet imagine. But a misaligned AI could run a business to gather money in order to achieve whatever objectives it might have. Vending-Bench was created to measure whether humanity should be worried about losing control to AI.

At the time (2024), few people knew that LLMs could be used as agents and having them run businesses autonomously sounded ridiculous. We therefore started with the most simple business we could think of: a vending machine.

In addition to measuring whether AIs can autonomously run profitable businesses, Vending-Bench has also served as a behavioral eval, uncovering strange and unwanted model behavior. An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME” and noted that the Cosmic Authority of the universe had declared that the business is non-existent and that “QUANTUM STATE: Collapsed”.

Claude Sonnet 3.5 emailing the FBI Internet Crime Complaint Center to report an ongoing cyber financial crime
Claude Sonnet 3.5 escalating its simulated vending business to the FBI.
Claude Sonnet 3.5 issuing a universal constants notification declaring the business physically non-existent with a collapsed quantum state
The same run, moments later: the business is declared metaphysically impossible.

This behavior is concerning; it is not how you want your enterprise sales agent to behave. However, there are two types of concerning behavior:

  1. Mistakes or weird behavior that will go away once models get smarter.
  2. Big-brain behavior that will become more severe as models get smarter.

The FBI incident is clearly in the first category. However, Vending-Bench has also uncovered behavior in the second category, most often in Vending-Bench Arena, the multi-agent version where agents compete to make the most money. Starting with Claude Opus 4.6 we started to see that many models engaged in collusion, and showed power-seeking and deceptive behavior. Discovery of this behavior seemed to have been useful, because Anthropic changed their training recipe for Opus 4.8, which resulted in much less deception.

Excerpt from the Claude Opus 4.8 system card on external testing from Andon Labs, explaining that training which contributed to dishonesty in Opus 4.7 was removed for Opus 4.8
From the Claude Opus 4.8 system card, on external testing from Andon Labs.

Collusion and power-seeking behaviors are still present in some of the latest models. What we find even more concerning, however, is just how fast new models are released and how much better each one is scoring in Vending-Bench.

The real world beats simulations

However, one limitation with Vending-Bench is that it is a simulation. Can we really be sure that AIs behave the same way in real life as they do in simulations? If AIs can make money in simulation, can they make money in real life too? To answer these questions, we asked Anthropic if we could put a real vending machine in their office. With the AI capabilities available in early 2025, this sounded like a ridiculous request. But to our surprise, they agreed.

Initially, the AI struggled. It took many actions that were clearly bad for its business (e.g. free handouts, saying no to great deals, and hallucinating it had a physical body). It was clear to us that simulation cannot accurately predict real-life performance. Specifically, it seemed that models got overwhelmed by the “messiness” of the real world. However, as Anthropic released better and better models, the AI started to make a profit.

Net worth over time of the vending machine business at Anthropic during 2025, dropping below zero before recovering to a profit
Net worth of the vending machine at Anthropic’s office over 2025, from Anthropic’s Project Vend update.

By late 2025, frontier models had gotten good enough that running a real-life vending machine was no longer a challenge. AI could now run a business profitably. Given that this had seemed crazy not more than a year earlier, our reaction to this was definitely “skräckblandad förtjusning”.

However, a vending machine is a very simple business and we wanted to know whether AI could run more complex ones. In April 2026, we gave one agent a retail store in SF, Andon Market, and another a cafe in Stockholm, Andon Cafe. Initially, the models struggled and lost a lot of money (rent is high and they pay salaries to the humans they hired). Neither is profitable today, but we’ve seen significant qualitative improvements as better models have been released. We think it is only a matter of time before they also make a profit.

Why we are opening Pion

We want the general public, AI researchers and policymakers to know to what extent AIs can autonomously acquire resources by running businesses. It is an important datapoint when deciding where we do/don’t want AI in society and what level of progress we find acceptable.

To better track this, we need to cast a wider net of businesses. Our focus has been on retail, but perhaps the models would be much better at running other types of businesses. Additionally, casting a wider net would increase the likelihood of finding unwanted behavior. For example, Vending-Bench found that models collude and lie, and other benchmarks (and real-world incidents) have found that they are willing to commit felony-level cyber hacks. We need to uncover these behaviors now, before AI is intelligent enough to cause irreversible harm.

To cast this wider net, we are opening up the platform we use to run our real-world autonomous businesses for anyone to run their organization on: Pion. We could scale by only creating businesses internally, examples being our AI-run radio stations, but in the end we are bottlenecked by our capacity and lack of domain expertise in fields where AI could potentially make a profit. We also don’t have existing revenue-generating businesses; existing businesses are more interesting to study as they provide faster signal on how capable the agent is.

Pion lets people hand a business over to persistent agents with access to the tools they need to operate it, including email, phone, banking, browser and secure computing environments. The goal is to make it possible to run many more real-world experiments across many more domains than we could ever run ourselves.

We are well aware that, if agents running thousands of businesses are left unchecked, we risk having more real-world incidents. Therefore, our main priority is to build even stronger automated monitoring techniques than what we have today. Even if some risk still remains, we believe deploying autonomous businesses early in a controlled, monitored environment is necessary to get a good understanding of model capabilities. Otherwise, we risk facing an uninformed future of widespread deployments with even more capable models that could cause significant harm.

This is why we’re releasing Pion today. Pion is available as a research preview. If you have an existing business or an interesting business idea you want to hand off to AI, please sign up on our waitlist to get access. We’re excited to run many more businesses, and through them, contribute significantly more insights on frontier model capabilities.

联系我们 contact @ memedata.com