1.4万份报告,3小时40分钟,四家一起趴下。
9月3日上午(美东时间),ChatGPT、Claude、Grok、Gemini——人们日常最常用的四个闭源AI服务——罕见地在同一时段中断。Ars Technica的形容是"practically unheard of"(几乎闻所未闻)。对普通用户这只是一次闹脾气,对把整个工作流押在这四家身上的团队,这是一场没有预案的压力测试。
时间线:从9:23到12:55
Anthropic最先出事,也最先报告:9:23 AM ET,状态页显示Claude Mythos 5.1、Fable 5.1与Opus 5部分中断,Sonnet 5出现短暂异常;12:16 PM恢复。OpenAI在10:43跟进,报告ChatGPT与Codex错误率升高,12:55恢复。Grok的曲线更陡:Downdetector报告数从9点的不及10份,45分钟内冲到1365份;Gemini从10:30的23份涨到11点后的412份,StatusGator显示其API在10:45至11:15间处于疑似中断。中文监测口径下,Downdetector的OpenAI相关报告超过1.2万份,Claude约1200份,Grok约1000份,合计超1.4万份——Cursor等下游编程工具也随之连坐。
根因各说各话,推测无人认领
直到全部恢复,也没有一份联合解释。OpenAI的说法是路由错误(routing error);Anthropic的说法是基础设施问题——两家口径不同,也都没有指向任何第三方。外界注意到,AWS、Azure、Cloudflare当天均未报告重大故障,于是"四家共用底层基础设施"的推测在社交平台流传;但目前没有任何一方确认共同根因,这个说法只能停在推测。值得记下的背景是:Claude过去90天的正常运行率为99.4%,ChatGPT为99.63%——这次"几乎闻所未闻",恰恰发生在各家可靠性最好的时期。
真正的考题:集中度
这次事故真正的信息量不在故障本身,而在故障的形状:不是一家出事,而是四家同时出事。过去两年,把工作流押进闭源模型的企业与开发者只增不减,"模型即供应商"的集中度一路抬升;当最头部的四家在同一分钟失联,"换一家"这个从容的备选项第一次集体失效。宕机总会发生,真正的问题是:当它以"集体"的形式发生时,你的降级方案里还剩下谁?
答案应该在事故前就写好——3小时40分钟,足够每个团队补一课。
明天见。
14,000 reports, 3 hours 40 minutes, four outages at once.
On the morning of September 3 (US Eastern time), ChatGPT, Claude, Grok and Gemini — the closed AI services people reach for daily — went down in a rare simultaneous window. Ars Technica called it "practically unheard of." To ordinary users it was a tantrum; to teams with entire workflows riding on these four, it was a stress test nobody had a playbook for.
Timeline: From 9:23 to 12:55
Anthropic went first and reported first: at 9:23 AM ET, its status page showed partial outages for Claude Mythos 5.1, Fable 5.1 and Opus 5, with a brief Sonnet 5 anomaly; recovery came at 12:16 PM. OpenAI followed at 10:43, reporting elevated errors on ChatGPT and Codex, restored by 12:55. Grok's curve was steeper: Downdetector reports jumped from under 10 at 9 AM to 1,365 by 9:45; Gemini climbed from 23 at 10:30 to 412 past 11 AM, with StatusGator flagging its API as likely down from 10:45 to 11:15. In Chinese-language tallies, Downdetector logged 12,000+ OpenAI-related reports, roughly 1,200 for Claude and about 1,000 for Grok — 14,000+ combined, with downstream tools like Cursor dragged down too.
Root Causes, Told Separately
Even after full recovery, there was no joint explanation. OpenAI's account: a routing error. Anthropic's: an infrastructure issue — different stories, neither pointing at a third party. Observers noted that AWS, Azure and Cloudflare reported no major incidents that day, so a theory about "shared underlying infrastructure" spread across social platforms; with no party confirming a common root cause, it remains exactly that — a theory. Worth pinning down: Claude's uptime over the past 90 days was 99.4%, ChatGPT's 99.63% — the "practically unheard of" happened precisely during everyone's most reliable stretch.
The Real Exam: Concentration
The real payload of this incident is not the outage but its shape: not one provider failing — four failing together. Over two years, enterprises and developers have piled ever more workflows onto closed models, and "model as vendor" concentration has only risen; when the top four go dark in the same minute, the calm fallback of "just switch" collectively fails for the first time. Outages will keep happening. The question is: when they arrive in plural form, who is left in your degradation plan?
The answer should have been written before the incident — 3 hours 40 minutes is enough for every team to make up the lesson.
See you tomorrow.
当四家最头部的能力供应商在同一分钟失联,"换一家"这个选项第一次集体失效。
—— Dawn Vision编辑部
When all four top capability providers go dark in the same minute, "just switch to another one" fails for the first time.
— The Dawn Vision Editorial Desk
9月3日美东上午四大闭源服务罕见重叠宕机 · Anthropic 9:23 AM ET最先报告(Claude Mythos 5.1/Fable 5.1/Opus 5部分中断,Sonnet 5短暂异常),12:16恢复 · OpenAI 10:43报告ChatGPT与Codex elevated errors,12:55恢复 · Grok Downdetector从<10(9AM)涨至1365(9:45);Gemini从23(10:30)涨至412(11AM+),StatusGator显示Gemini API 10:45-11:15 likely outage · 中文口径:OpenAI超1.2万份、Claude约1200、Grok约1000,合计超1.4万份;全程约3小时40分钟 · Cursor等下游工具连带 · 根因:OpenAI 'routing error',Anthropic 'infrastructure issue';AWS/Azure/Cloudflare均无重大故障报告;共享基础设施推测无确认 · 可靠性背景:Claude 90天99.4%,ChatGPT 99.63%
Four major closed services went down in a rare overlapping window on Sept 3 (US Eastern) · Anthropic reported first at 9:23 AM ET (Claude Mythos 5.1 / Fable 5.1 / Opus 5 partial outage; brief Sonnet 5 anomaly), recovered 12:16 · OpenAI at 10:43 reported elevated ChatGPT and Codex errors, recovered 12:55 · Grok: Downdetector <10 at 9 AM to 1,365 at 9:45; Gemini: 23 at 10:30 to 412 past 11 AM, StatusGator likely-outage 10:45-11:15 · Chinese tallies: 12,000+ OpenAI, ~1,200 Claude, ~1,000 Grok, 14,000+ combined; ~3h40m total · Downstream tools like Cursor affected · Root causes: OpenAI 'routing error', Anthropic 'infrastructure issue'; no major AWS/Azure/Cloudflare incidents; shared-infrastructure theory unconfirmed · Reliability context: Claude 99.4% over 90 days, ChatGPT 99.63%
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理当日采集信号并交叉验证 2 个核心来源后生成,经编辑部人工审核。素材来源:Ars Technica(时间线与"几乎闻所未闻"定性)、The Verge(各服务状态页口径);Downdetector中文监测数据来自中文媒体转述口径。各官方根因说法不同,正文未合并为单一结论。
This article was generated by the Dawn Vision cognitive engine processing collected daily signals with cross-validation of two core sources, followed by human editorial review. Sources: Ars Technica (timeline and the 'practically unheard of' framing), The Verge (status-page accounts); Chinese-language Downdetector tallies come from Chinese media relay. Official root-cause accounts differ and were not merged into a single conclusion.