9月12日,硅谷最大的竞争对手们做了一件罕见的事:同一天,站在了同一边。
Anthropic CEO Dario Amodei在公司官网发布了一篇超过5000字的长文——"We Must Pace the Frontier"。这不是一篇学术论文,不是一篇愿景声明,而是一份带着具体执行方案的行动纲领。几个小时内,OpenAI CEO Sam Altman公开表示支持,xAI(已被SpaceX收购后的SpaceXAI)和Microsoft也相继表态。同日,OpenAI宣布2026年不会进行IPO,并将这一决定与安全承诺直接挂钩。
一天之内,前沿AI安全从"大家觉得应该做"变成了"有人已经开始做了"。这个转变值得仔细拆解。
三步计划:从纸面到工牌
Amodei提出的三步计划,每一步都比前一步更具体、更难、也更重要。
第一步:嵌入式评估员(Embedded Evaluators)。这是整个计划中最具操作性的部分。Amodei提议,AI实验室应该允许独立第三方评估员获得与员工相同级别的访问权限——包括工牌、笔记本电脑、内部系统权限。这些评估员可以在实验室内部自由活动、审查安全流程、测试模型行为,并且有权随时公开发布自己的发现。
这与现有的"红队测试"有本质区别。红队测试通常是限时、限定范围的项目,测试结束即终止。嵌入式评估员则是一种长期的、持续的、有独立性的监督机制。他们的存在本身就是一种威慑——如果一家AI实验室知道每天都有独立评估员在看它的代码和决策,它的行为自然会发生变化。
Anthropic承诺率先单方面接受嵌入式评估员,这意味着不等行业共识形成,Anthropic就先在自己身上做实验。这是一个聪明的策略:如果实验成功,其他公司如果不跟进,就面临巨大的舆论压力。
第二步:民主国家行业协调。Amodei提议,AI前沿实验室之间应该建立更紧密的安全信息共享和协调机制——但这只在"民主国家"的实验室之间进行。这一步的核心障碍是美国反垄断法。竞争对手之间共享安全信息、协调政策,很可能被解读为限制竞争的"合谋"。因此,Amodei明确呼吁:国会需要为AI安全协调提供反垄断豁免。
这是一个大胆但必要的主张。如果没有反垄断豁免,前沿实验室之间的安全协调就永远停留在"自愿"层面,缺乏强制性和可执行性。Amodei等于在说:如果你们真的在乎AI安全,那就给法律开一道口子。
第三步:全球协调(含中国)。这是最具争议也最具远见的一步。Amodei承认,在AI安全问题上排除中国是不现实的——中国在AI研究上的实力决定了,任何不包含中国的全球安全框架都是不完整的。他提议探索一种包括中国在内的全球协调机制,即使这在当前的地缘政治环境下极其困难。
OpenAI的响应:不上市是安全承诺的货币化
同一天,OpenAI的回应同样引人注目。
Sam Altman在社交媒体上公开支持Amodei的三步计划,称其为"建设性的前进方向"。但更具实质意义的是OpenAI同步宣布的两项决定:
第一,OpenAI宣布2026年不进行IPO。表面上看,这是对"上市时间表"的推迟。但将这一声明放在Amodei的三步计划的语境下,含义就不同了。IPO意味着利润最大化压力、季度财报压力、股东利益优先。一个正在进行IPO的公司很难在安全问题上做出大胆的决策——因为安全投入是成本,短期内看不到回报。OpenAI选择不上市,等于给自己保留了在安全问题上"做正确的事"的空间。
第二,OpenAI同日发布了"针对前沿AI的安全要求框架"(A Framework for Safety Requirements for Frontier AI)。这份文件试图定义:什么样的AI模型需要接受什么样的安全审查?门槛在哪里?谁来执行?虽然具体细节还有待完善,但作为对Amodei呼吁的实质性回应,它比一句"我们支持安全"有意义得多。
"嵌入式评估员不是审计师,不是顾问——他们是拿工牌的独立观察者,有权随时公开发布发现。这才是真正的透明。"—— Dawn Vision编辑部
为什么这次不一样:从自愿到承诺
AI安全领域从来不缺漂亮话。每隔几个月就有一封"联名信"、一份"承诺书"、一场"峰会"。但这次有几个关键区别。
区别一:Anthropic率先单方面行动。以往的安全倡议都是"大家一起做"——这意味着没有人为失败负责。Anthropic承诺自己先做嵌入式评估员,这是一个可验证、可追踪的承诺。如果Anthropic真的引入了独立评估员并允许他们公开报告,这将成为整个行业的标杆。如果Anthropic食言,也会被迅速发现。
区别二:OpenAI用不上市做筹码。IPO对OpenAI来说是巨大的诱惑——数十亿美元的回报。OpenAI选择推迟IPO并将之与安全挂钩,意味着它在用真金白银为安全承诺背书。这不是口头支持,这是资本的机会成本。
区别三:计划包含了具体机制。嵌入式评估员有工牌、有笔记本、有权限、可以公开发表——这是具体的。反垄断豁免需要国会行动——这是具体的。全球协调包含中国——这需要外交突破,但至少定义了目标。与以往那些"我们将致力于安全"的模糊承诺相比,这份计划每个步骤都有可操作的路径。
挑战与未来:三步计划能走多远?
即使在最好的情况下,Amodei的三步计划也面临巨大的执行挑战。
嵌入式评估员制度面临的第一个问题是资金来源。谁来支付评估员的薪水?如果由AI实验室支付,评估员的独立性如何保证?如果由政府或独立基金支付,资金从哪里来?Amodei在文中没有给出明确答案。
第二个问题是评估标准。评估员"有权公开发布发现"——但"发现"的标准是什么?什么样的发现构成"安全风险"?谁来定义这些标准?如果没有统一的标准,评估员的报告可能沦为一场各说各话的PR游戏。
反垄断豁免更是充满了不确定性。美国国会目前在任何AI相关立法上都难以达成共识,更不用说为竞争对手之间的"合谋"开绿灯。即使在两党都支持AI监管的大框架下,反垄断豁免也会遭到许多议员的反对——因为它触及了美国竞争法的核心原则。
全球协调(含中国)的难度更是不言而喻。中美在AI领域的竞争日益激烈,双方在AI安全标准上存在根本分歧。让两国在AI安全问题上达成实质性协议,可能需要比Amodei想象的更长的时间。
但即使如此,这份计划的价值不在于它是否能全部实现,而在于它定义了一个方向。过去的AI安全讨论是抽象的——"我们应该怎么做"、"我们需要什么原则"。Amodei把讨论拉到了具体的机制层面——"谁来做"、"怎么做"、"做到什么程度"。即使最终的实施方案与Amodei的提议不同,这份计划也为后续的讨论设定了基准。
9月12日可能不会被铭记为"AI安全胜利日"。但它很可能会被视为AI安全从口号走向执行的转折点。当Anthropic率先让独立评估员走进办公室,当OpenAI用推迟IPO来表达对安全的承诺,当行业竞争对手开始讨论如何协调而非各自为战——这些细节比任何宣言都更有说服力。
AI安全的执行时代,才刚刚开始。
明天见。
On September 12, Silicon Valley's fiercest rivals did something rare: on the same day, they stood on the same side.
Anthropic CEO Dario Amodei published a 5,000+ word essay on the company website — "We Must Pace the Frontier." This isn't a research paper or a vision statement. It's an action plan with specific execution blueprints. Within hours, OpenAI CEO Sam Altman publicly endorsed it, xAI (now SpaceXAI after SpaceX's acquisition) and Microsoft followed suit. The same day, OpenAI announced it will not IPO in 2026, explicitly tying the decision to its safety commitments.
In a single day, frontier AI safety went from "everyone agrees we should do something" to "someone has actually started doing something." That shift deserves a close look.
Three Steps: From Paper to Badges
Each step of Amodei's three-step plan is more concrete, harder, and more important than the last.
Step One: Embedded Evaluators. This is the most operationally concrete piece of the entire plan. Amodei proposes that AI labs grant independent third-party evaluators employee-level access — including badges, laptops, internal system permissions. These evaluators can move freely inside the lab, audit safety processes, probe model behavior, and publish their findings publicly at any time.
This is fundamentally different from existing "red-teaming." Red-team exercises are typically time-bound, scope-limited projects that end when the test ends. Embedded evaluators represent a long-term, continuous, independent oversight mechanism. Their very presence is a deterrent — if a lab knows independent evaluators are watching its code and decisions daily, behavior changes on its own.
Anthropic committed to unilaterally accepting embedded evaluators first, meaning it won't wait for industry consensus — it'll experiment on itself. This is a clever strategy: if the experiment succeeds, other companies face enormous public pressure to follow. If they don't, the optics are devastating.
Step Two: Democratic Industry Coordination. Amodei proposes tighter safety information-sharing and coordination mechanisms among frontier AI labs — but only among labs in "democratic nations." The core obstacle here is U.S. antitrust law. Competitors sharing safety information and coordinating policy could easily be construed as anticompetitive "collusion." Amodei explicitly calls on Congress to provide antitrust exemptions for AI safety coordination.
It's a bold but necessary demand. Without antitrust exemptions, safety coordination among frontier labs remains "voluntary" — toothless and unenforceable. Amodei is essentially saying: if you truly care about AI safety, carve out a legal exception.
Step Three: Global Coordination (Including China). This is the most controversial and most visionary step. Amodei acknowledges that excluding China from AI safety discussions is unrealistic — China's AI research capabilities mean any global safety framework without China is incomplete. He proposes exploring a coordination mechanism that includes China, even though this is extraordinarily difficult in the current geopolitical climate.
OpenAI's Response: Not IPOing Is Safety Commitment Monetized
OpenAI's response on the same day was equally striking.
Sam Altman publicly endorsed Amodei's three-step plan on social media, calling it a "constructive path forward." But the substantive moves were two simultaneous announcements from OpenAI:
First, OpenAI declared it will not IPO in 2026. On the surface, this is a delay. But in the context of Amodei's plan, the meaning shifts. An IPO means profit-maximization pressure, quarterly earnings pressure, shareholder primacy. A company mid-IPO can't make bold safety decisions — because safety spending is a cost with no short-term payoff. By not going public, OpenAI preserves room to "do the right thing" on safety.
Second, OpenAI simultaneously released "A Framework for Safety Requirements for Frontier AI". This document attempts to define: what kind of AI models need what kind of safety review? Where are the thresholds? Who enforces? Details still need fleshing out, but as a substantive response to Amodei's call, it's worth infinitely more than "we support safety."
"Embedded evaluators aren't auditors or consultants — they're independent observers with badges, empowered to publish findings at any time. That's real transparency."— The Dawn Vision Editorial Desk
Why This Time Is Different: From Voluntary to Committed
The AI safety field has never lacked pretty words. Every few months brings another "open letter," "pledge," or "summit." But this time has key differences.
Difference One: Anthropic is acting unilaterally first. Past safety initiatives were always "let's all do this together" — meaning no one was responsible for failure. Anthropic committing to embedded evaluators first is a verifiable, trackable promise. If Anthropic actually brings in independent evaluators and allows public reporting, it becomes the industry benchmark. If Anthropic backtracks, that too will be swiftly exposed.
Difference Two: OpenAI is using its IPO as collateral. The IPO is an enormous temptation for OpenAI — billions in returns. OpenAI choosing to delay and tie it to safety means it's backing its safety commitment with real capital opportunity cost. This isn't vocal support; it's financial skin in the game.
Difference Three: The plan includes specific mechanisms. Embedded evaluators get badges, laptops, access, and publish rights — that's concrete. Antitrust exemptions require Congressional action — that's concrete. Global coordination includes China — that requires diplomatic breakthroughs but at least defines a target. Compared to the vague "we will commit to safety" pledges of the past, every step in this plan has an actionable pathway.
Challenges Ahead: How Far Can This Plan Go?
Even under the best circumstances, Amodei's three-step plan faces massive execution challenges.
The embedded evaluator program's first problem is funding. Who pays the evaluators? If labs pay, how is independence guaranteed? If government or an independent fund pays, where does the money come from? Amodei doesn't provide clear answers.
The second problem is evaluation standards. Evaluators can "publish findings" — but what counts as a "finding"? What constitutes a "safety risk"? Who defines these standards? Without unified criteria, evaluator reports risk devolving into a PR game of dueling narratives.
Antitrust exemptions face enormous uncertainty. Congress can barely reach consensus on any AI-related legislation, let alone greenlight "collusion" among competitors. Even under a broad framework where both parties support AI regulation, antitrust exemptions will face fierce opposition — they strike at the core principles of American competition law.
The difficulty of global coordination (including China) is self-evident. U.S.-China AI competition intensifies daily, with fundamental disagreements on AI safety standards. Reaching substantive agreements between the two nations on AI safety may take far longer than Amodei imagines.
But even so, the plan's value lies not in whether it's fully implemented, but in the direction it defines. Past AI safety discussions were abstract — "what should we do" and "what principles do we need." Amodei pulled the discussion into specific mechanisms — "who does it," "how it's done," and "to what degree." Even if the final implementation differs from Amodei's proposal, this plan sets the baseline for all future discussion.
September 12 may not be remembered as "AI Safety Victory Day." But it may well be seen as the inflection point when AI safety moved from slogans to execution. When Anthropic first lets independent evaluators walk through the door, when OpenAI uses an IPO delay to express safety commitment, when industry competitors start discussing coordination instead of going it alone — these details are more persuasive than any manifesto.
The execution era of AI safety has only just begun.
See you tomorrow.
Dawn Vision, Dario Amodei, Anthropic, AI safety, frontier AI, embedded evaluator, OpenAI, xAI, antitrust exemption