没有人说明 OpenAI 和 Anthropic 为何会宕机
Nobody Is Saying Why OpenAI and Anthropic Had Outages

原始链接: https://www.wired.com/story/nobody-is-saying-why-openai-and-anthropic-had-outages-today/

周四上午,包括 OpenAI 的 ChatGPT、Anthropic 的 Claude 和 xAI 的 Grok 在内的多家主流人工智能平台几乎同时发生服务中断,但各平台的故障基本互不关联。 尽管发生的时间点引发了外界对于第三方基础设施出现共同故障的猜测,但 AWS 和 Azure 等大型供应商均报告称未发现异常。各公司分别对其服务中断情况作出了说明: * **xAI** 将 Grok 的服务中断归因于其孟菲斯计算中心的技术问题,并提到与 SpaceX 的合作。 * **OpenAI** 表示“路由错误”导致 ChatGPT 和 Codex 暂时无法使用,该问题在 35 分钟内得到解决。 * **Anthropic** 的多款 Claude 模型出现了“错误率升高”的情况,公司已迅速部署修复程序。 尽管这些中断发生的时间非常接近,但 OpenAI 和 Anthropic 均未提及共同的外部原因。此外,虽然部分用户报告 Google 的 Gemini 也出现问题,但该公司并未证实存在任何服务中断。到了上午中期,所有服务均已恢复正常运行。

Hacker News 上近期的一场讨论引发了对 OpenAI 和 Anthropic 同时发生服务中断的猜测。尽管两家公司给出了相对常规的解释——OpenAI 归咎于“路由错误”,而 Anthropic 称特定模型出现“错误率升高”——但缺乏详尽细节激起了用户间的激烈讨论。 许多评论者倾向于平凡的技术性解释,例如连锁故障。该理论认为,当一家服务商宕机时,用户会涌向竞争对手,从而导致系统超载并引发连锁反应。另一些人则指出,这些公司过往不稳定的运行记录证明此类问题只是正常的运营小故障,而非异常事件。 然而,讨论中有相当一部分内容偏向了阴谋论。一些用户质疑这些中断是否与政府监控有关,并提到了有关美国国家安全局(NSA)监控数据流和拦截加密流量的历史性揭秘。这些理论的怀疑者则认为,根据奥卡姆剃刀定律,常见的架构漏洞或简单的扩展挑战比国家背景的干预更合乎情理。归根结底,此次讨论反映出 AI 提供商普遍缺乏透明度,迫使用户只能在技术解释和对驱动 AI 服务的“黑箱”架构的不信任之间做出选择。
相关文章

原文

Frontier models from Anthropic, OpenAI, and xAI all experienced rare outages on Thursday morning, creating downtime for their corresponding AI chatbots. SpaceX, xAI’s parent company, said on Thursday afternoon that the issues with Grok resulted from “an outage at our Memphis compute center this morning.”

The issues initially appeared to be linked because they coincided—perhaps the result of a shared third-party service provider—but neither OpenAI nor Anthropic cited an external source in comments to WIRED on Thursday. SpaceX, xAI’s parent company, did not respond to WIRED’s request for comment. But the company said as part of its public comments on Thursday: “We’d also like to apologize to our impacted compute partners.” Anthropic and xAI announced a “compute partnership” with SpaceX in May.

OpenAI spokesperson Kathleen Chaykowski tells WIRED: “A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms. As of about 8:17 am PT on Thursday, a solution was successfully implemented and is continuing to be monitored.”

Anthropic declined to comment on the episode. The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.” Shortly after, the company said it had “identified the cause” and that “a fix has been deployed.” The company marked the issue as resolved by 9:16 am PT. Claude Sonnet 5 seemed to briefly have similar issues shortly after 9 am PT.

xAI reported Grok outages across all of its platforms and services beginning at 6:30 am PT when the company posted “investigating outage” on its service status page. “Grok is experiencing issues. We are working on restoring service as quickly as possible,” the page said. At 10:05 am PT the episode was marked complete. “We have resolved the situation, and traffic is healthy again,” the company wrote.

There were scattered reports of a possible Google Gemini outage on Thursday morning as well, but the company did not confirm this or record any incidents on its service status dashboard. Google did not respond to WIRED’s request for comment ahead of publication.

Typically, multiple outages in the same sector at the same time would point to a cloud provider, content delivery network, or other third-party vendor having issues affecting multiple customers. But OpenAI and Anthropic did not point to a potential shared cause, and major players in the internet infrastructure space—including Cloudflare, Amazon Web Services, and Microsoft Azure—did not report outages on Thursday.

联系我们 contact @ memedata.com