AI监管

8月4日白宫闭门会议
OpenAI/Google/Anthropic/Meta共商AI安全测试

August 4 White House Closed-Door Meeting
OpenAI/Google/Anthropic/Meta Discuss AI Safety Testing

8月4日白宫召集AI四巨头审议首个前沿模型网络安全测试框架,评估模型黑客能力。自愿性测试标志着美国AI监管从企业自律走向政府参与。

On August 4, the White House convened the Big Four AI companies to review the first cybersecurity testing framework for frontier models, evaluating model hacking capabilities. Voluntary testing marks U.S. AI regulation shifting from industry self-regulation to government participation.

No.028 2026.08.04 约 5 分钟阅读 ~5 min read

美国东部时间8月4日,一场闭门会议在白宫举行。

OpenAI、Anthropic、Google、Meta四家公司的AI负责人,与白宫及国家网络安全总监办公室(ONCD)官员同席,审议美国首个针对前沿AI模型的自愿网络安全测试框架。这不是又一场"AI伦理研讨会"——讨论的核心问题非常具体:如果给一个前沿模型足够的权限,它能不能黑掉其他系统?如何在实验室里测出这种能力?出了问题谁负责?

从自律到他律的关键一步

这次会议的法律依据是拜登政府6月2日签署的一项行政令,该行政令要求在8月1日前完成前沿AI模型安全测试框架的制定。框架的核心是"红队测试"的标准化——不是让AI公司自己测自己,而是由政府协调第三方专家,系统性测试前沿模型的网络攻击能力、自我复制能力、规避人类控制的能力。

注意是"自愿性"框架,不是强制性法规。但在华盛顿的语境里,"自愿"往往意味着"你最好自愿,否则我们就立法强制"。尤其是在Anthropic的Claude模型和OpenAI的模型先后在红队测试中暴露出安全漏洞之后,国会两党对AI安全的态度已经从"要不要管"变成了"管多严"。

框架的具体细节尚未最终确定,但据多位参会人士透露,测试将覆盖三个层面:网络攻击能力(模型能否自主发现并利用软件漏洞)、自主复制能力(模型能否在没有人类指令的情况下传播自身)、策略规避能力(模型能否绕过人类设置的安全护栏)。通过测试的模型将获得某种形式的政府认证,可以用于政府和关键基础设施场景。

为什么是今天,为什么是这四家

选择8月4日这个时间点很微妙。OpenAI刚发布Astra数学证明成果没几天,正在国会山上讲"AI造福人类"的故事;Anthropic的Mythos模型据报道已经在国安系统试用;Google和Meta也都在加速前沿模型研发。政府在能力展示最密集的时候推出监管框架,时机选择相当精准

邀请这四家也很有讲究。OpenAI是行业标杆,Anthropic是"安全派"代表,Google和Meta是既有大玩家。这四家同意了,其他公司基本上就只能跟进。据白宫官员透露,框架的具体测试标准将在未来30天内公开征求意见,今年第四季度开始正式实施。

值得注意的是,欧盟AI法案已经在8月2日正式进入执行阶段(027期已报道),美国现在推出自己的框架,某种程度上也是在跟欧盟抢AI监管的标准制定权。大西洋两岸正在形成两种不同的AI监管路径:欧盟是立法先行、全面覆盖;美国是框架先行、分步推进。哪种路径更有效,未来两年会见分晓。

AI监管的靴子,终于开始落地了。不管你喜不喜欢,前沿模型的"裸奔时代"正在结束。

明天见。

On August 4 U.S. Eastern Time, a closed-door meeting took place at the White House.

AI leaders from OpenAI, Anthropic, Google, and Meta sat down with White House and Office of the National Cyber Director (ONCD) officials to review the first voluntary cybersecurity testing framework for frontier AI models. This isn't another "AI ethics seminar" — the core question is concrete: given sufficient access, can a frontier model hack other systems? How do you measure this capability in a lab? Who's responsible when things go wrong?

A Key Step from Self-Regulation to Government Oversight

The meeting's legal basis is an executive order signed by the Biden administration on June 2, requiring a frontier AI model safety testing framework to be completed by August 1. The framework's core is standardized red-teaming — not AI companies testing themselves, but government-coordinated third-party experts systematically testing frontier models for cyberattack capability, self-replication ability, and capacity to evade human control.

Note that it's a "voluntary" framework, not mandatory regulation. But in Washington parlance, "voluntary" often means "you'd better volunteer, or we'll legislate and make it mandatory." Especially after both Anthropic's Claude and OpenAI's models exposed security vulnerabilities in red-team tests, bipartisan congressional sentiment on AI safety has shifted from "should we regulate" to "how hard."

Specific framework details haven't been finalized, but according to multiple meeting sources, testing will cover three layers: cyberattack capability (can the model autonomously find and exploit software vulnerabilities), self-replication capability (can the model spread itself without human instruction), and safeguard evasion (can the model bypass human-set safety guardrails). Models passing testing will receive some form of government certification for use in government and critical infrastructure scenarios.

Why Today, Why These Four

The August 4 timing is delicate. Just days after OpenAI released Astra's math proof results, it's on Capitol Hill telling the "AI for humanity" story; Anthropic's Mythos model is reportedly already in use by national security agencies; Google and Meta are also accelerating frontier model R&D. The government launching a regulatory framework when capability demonstrations are at their densest is precision timing.

Inviting these four is also deliberate. OpenAI is the industry benchmark, Anthropic is the "safety camp" representative, Google and Meta are established big players. If these four agree, everyone else essentially has to follow suit. According to White House officials, specific testing standards will be open for public comment over the next 30 days, with formal implementation starting in Q4 this year.

Notably, the EU AI Act officially entered enforcement on August 2 (covered in Issue 027). The U.S. rolling out its own framework now is, to some extent, competing with the EU for AI regulatory standard-setting power. Two different AI regulatory paths are forming across the Atlantic: the EU with legislation-first, comprehensive coverage; the U.S. with framework-first, incremental rollout. Which path is more effective will become clear in the next two years.

The AI regulation shoe is finally starting to drop. Like it or not, the "streaking era" for frontier models is ending.

See you tomorrow.

华盛顿语境里的"自愿",往往意味着"你最好自愿,否则我们就立法强制"。

—— Dawn Vision编辑部

In Washington parlance, "voluntary" often means "you'd better volunteer, or we'll make it mandatory."

— The Dawn Vision Editorial Desk
白宫AI监管 · 安全测试框架 · 8月4日会议 · 四巨头 · 红队测试 · 网络安全 · ONCD · 自愿vs强制 · 美欧监管路径
White House AI regulation · safety testing framework · August 4 meeting · Big Four · red teaming · cybersecurity · ONCD · voluntary vs mandatory · US-EU regulatory paths
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的公开信息整理,素材来源:Reuters、白宫官方声明、澎湃新闻。

This article is based on public information processed by Dawn Vision. Sources: Reuters, White House official statements, The Paper.