Focus · 焦点

Anthropic 15亿版权和解落地
AI训练数据边界终划定

Anthropic $1.5B Settlement Landed
AI Training Data Boundary Set

美国版权史最大金额和解——15亿美元、50万部作品、每部3000美元。法官裁定AI训练属合理使用,但盗版获取书籍违法。一个字:买的可以,偷的不行。

The largest settlement in U.S. copyright history — $1.5 billion, 500,000 works, $3,000 per work. The judge ruled AI training is fair use, but pirating books to train on is illegal. Simple: bought is fine, stolen isn't.

No.018 2026.07.21 约 10 分钟阅读 ~10 min read

15亿美元。

7月20日,美国旧金山联邦法院正式批准了Anthropic与作家群体之间的版权诉讼和解——这是美国版权诉讼史上金额最高的集体和解。50万部受版权保护的作品,每部作品获得约3000美元赔偿,总额15亿美元(约合102亿元人民币)。签字的法官从Alsup换成了Martinez-Olguin,但判决的核心逻辑没变:AI训练本身是合理使用,但你获取训练数据的方式如果是盗版,那就是违法。

一句话总结:买的书可以用来训练,偷的不行。

这个看似简单的结论,却可能重塑整个AI产业的数据供应链。过去几年,所有大模型公司都在同一个灰色地带狂奔——爬取全网内容、下载盗版电子书、抓取社交媒体数据,没人知道法律的红线在哪里。现在,第一根红线划下来了。

判决的核心:合理使用,但获取方式非法

这个案子最值得玩味的地方在于:原告和被告各赢了一半。

Alsup法官在去年的初步裁决中,在最核心的问题上站在了Anthropic一边——他裁定用受版权保护的文本来训练AI模型属于合理使用(fair use)。这是AI产业最想要的答案:训练本身不侵权。如果训练就是侵权,那全球所有大模型公司的基础都不成立,OpenAI、Google、Meta、Anthropic全都得吃官司吃到破产。

但Anthropic高兴得太早了。Alsup法官同时认定,Anthropic获取这些书的方式有问题——它的训练语料来自两个渠道:一是合法购买并扫描的书籍,二是从Library Genesis、Pirate Library Mirror等盗版网站下载的书籍。第二个渠道,本身就是违法的。

这就像什么呢?就像你去图书馆借了本书,读完之后把内容记在脑子里,这不叫偷——这叫学习,合理使用。但如果你是去偷了一本书来读,那偷书这个行为本身就是犯法的,跟你读完之后用这些知识做什么没关系。

Alsup法官本来要把「盗版获取」这个问题交给陪审团审判,Anthropic赶紧同意和解——因为谁也不知道陪审团会判出多少赔偿金。15亿美元虽然是天价,但总比陪审团一怒之下判个50亿、100亿强。

所以这个判决的真正含义是:AI训练的合法性被确认了,但数据获取的合规成本被推高了。以后大模型公司不能再随便从盗版网站扒书了,要么花钱买授权,要么只用公开合法的数据源。

15亿美元到底贵不贵?

15亿美元,听起来是个天文数字。但算算账你会发现,这笔钱可能比想象中「便宜」。

50万部作品,每部3000美元。对于一本畅销书来说,3000美元可能连一年的版税都不到;但对于绝大多数普通作家来说,一本书一辈子可能都赚不到3000美元。很多小众作家、学术作者、诗集作者,他们的作品可能总共只卖了几百本,现在突然拿到3000美元的赔偿,相当于天上掉馅饼。

而对于Anthropic来说呢?15亿美元是什么概念?它最近一轮融资的估值是600多亿美元,15亿大约是估值的2.5%。换个角度看,Anthropic刚刚和TeraWulf签了20年190亿美元的算力合同——15亿的赔偿金,还不到一份算力合同的十分之一。

更重要的是,这个和解没有成为判例。因为是和解结案,不是上诉法院的判决,所以对其他案件没有约束力。Google、Meta、OpenAI还在打各自的版权官司,每个法官都可以有自己的判断。上周,Hachette、Cengage、Elsevier等大型出版商和作家Scott Turow等人才刚对Google提起集体诉讼,指控Google用他们的版权作品训练Gemini。

"训练本身是合理使用,但盗版获取违法——这个判决的精髓在于:它没有禁止AI,它只是要求AI公司像正常人一样付钱买东西。"—— Dawn Vision编辑部

所以15亿美元到底是贵是便宜,取决于你怎么看。从作家的角度看,每本书3000美元,比起被白嫖当然是强多了,但比起AI公司从这些内容中获得的价值,可能还是九牛一毛。从Anthropic的角度看,15亿买了一个「训练合法」的模糊确认,还避免了陪审团可能给出的天价赔偿,这笔交易其实不亏。

中国模型公司的意外利好

这个判决出来之后,最该偷笑的可能不是Anthropic,而是中国的大模型公司。

为什么?因为中国的大模型公司从一开始就面临更严格的数据合规要求。国内的监管环境决定了,中国大模型公司在训练数据的获取上一直相对「规矩」——不能随便爬,不能随便用,版权问题从一开始就是悬在头上的剑。结果就是,中国大模型公司在数据合规上的成本,反而可能比美国同行更低(因为一开始就按规矩来,不会积累巨大的法律风险)。

就在这个判决出来的同时,另一条新闻也值得关注:The Verge发表了一篇长文,标题就叫《中国对美国AI优势的组合拳》。文章说,月之暗面的Kimi K3和阿里的Qwen正在从两个方向夹击美国AI公司——一个用性能逼近GPT-5和Claude的闭源模型打高端,一个用完全开源的模型打生态,两者加起来,美国的AI领先优势正在快速收窄。

预测市场Polymarket上,Anthropic年底前估值达到1.5万亿美元的概率已经大幅下降,交易者把这个变化和Kimi K3的发布直接联系起来。与此同时,港股的智谱AI午后暴涨超过30%——公司落地了1GW国产算力中心,还完成了对中科加禾的收购。

版权判决的深层影响在这里:如果美国AI公司必须为训练数据支付更高的合规成本,而中国公司在数据合规上早已轻装上阵,那么两者之间的成本差距会进一步缩小。过去大家默认美国AI公司「跑得快」,但跑在前面的人身上挂的法律雷也更多。一旦这些雷陆续爆炸,后面的人反而可能追上来。

而且不要忘了,中国有全世界最大的中文语料库,也有独特的监管框架。国内的生成式AI管理办法早就出台了,数据合规的边界相对清晰。从这个角度看,先规范的市场反而可能先受益——就像先修路的地方先通车一样。

终局判断:数据成本时代到来

把15亿的版权和解、中国模型的崛起、MCP协议的演进放在一起看,一个清晰的产业拐点正在浮现:AI的竞争正在从「谁的模型大」转向「谁的数据干净、谁的成本可控、谁的生态完善」。

2023年比参数,2024年比性能,2025年比价格,2026年比什么?比合规、比数据质量、比基础设施。

这个判决会带来几个深远影响:

第一,高质量版权数据的价格会暴涨。以前出版社的内容被AI公司白嫖,现在有了15亿的先例,出版社腰杆子硬了,授权费肯定要涨。四大出版社(企鹅兰登、哈珀柯林斯、Simon & Schuster、Hachette)手里掌握的内容,会变成AI时代的「石油开采权」。以后大模型公司要做高质量模型,得先跟这些出版社谈授权——这又回到了「谁钱多谁说了算」的老路上。

第二,开源模型的相对优势会进一步扩大。闭源模型公司需要为训练数据支付巨额成本,但开源社区天然有更多合法数据来源——开发者自愿贡献的代码、开放获取的学术论文、Creative Commons授权的内容等等。如果闭源模型的成本因为版权问题持续上升,而开源模型的性能不断逼近,市场格局会发生什么变化?答案不言自明。

第三,数据合成会成为显学。既然真实世界的数据有版权问题,那我用AI生成的数据来训练AI行不行?这就是「合成数据」的思路。现在已经有很多研究在探索用大模型生成训练数据来训练下一代模型,这条路如果走通了,版权问题就不再是瓶颈——AI自己给自己造粮食。当然,合成数据也有自己的问题,比如模型坍缩、多样性不足,但比起动辄十几亿的版权费,这些技术问题反而是便宜的。

第四,全球AI监管会进入「拼细节」阶段。韩国的AI基本法今天(7月21日)正式生效,欧盟的AI法案7月1日全面实施,中国的生成式AI管理办法早已落地,美国还在打官司找边界。不同地区的监管框架不一样,AI公司的合规成本也不一样,这本身就会成为一种竞争力。哪里的监管清晰、可预期、成本合理,AI产业就会往哪里聚集。

15亿美元的和解不是终点,只是起点。它是AI产业从「野蛮生长」转向「合规发展」的第一个标志性事件。以后还会有更多的版权官司、更多的监管政策、更多的合规成本。AI公司不能再只盯着模型性能和融资估值了,数据供应链的合法性,正在变成下一个核心竞争力。

毕竟,在一个数据就是石油的时代,你首先得确保你用的石油——不是偷来的。

明天见。

$1.5 billion.

On July 20, a U.S. federal court in San Francisco formally approved the copyright settlement between Anthropic and a group of authors — the largest class-action copyright settlement in U.S. history. Half a million copyrighted works, roughly $3,000 per work, totaling $1.5 billion (~10.2 billion yuan). The judge signing off changed from Alsup to Martinez-Olguin, but the core logic remains: AI training itself is fair use, but pirating the data you train on is illegal.

One sentence sums it up: Bought books are fine to train on. Stolen ones aren't.

This seemingly simple conclusion could reshape the entire AI industry's data supply chain. Over the past few years, all LLM companies have been racing through the same gray area — scraping the entire web, downloading pirated e-books, hoovering up social media data, with no one knowing where the legal red line was. Now the first red line has been drawn.

The Core of the Ruling: Fair Use, but Illegal Acquisition

The most fascinating part of this case is that both sides won — sort of.

In his preliminary ruling last year, Judge Alsup sided with Anthropic on the most critical question: he ruled that training AI models on copyrighted text counts as fair use. This is the answer the AI industry wanted most: training itself isn't infringement. If training were infringement, the foundation of every major LLM company would collapse — OpenAI, Google, Meta, Anthropic would all get sued into bankruptcy.

But Anthropic celebrated too early. Judge Alsup also found that how Anthropic obtained those books was problematic — its training corpus came from two sources: books it legally purchased and scanned, and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. The second channel is illegal on its own terms.

Here's the analogy: if you borrow a book from the library, read it, and remember what you learned, that's not stealing — that's learning, fair use. But if you steal a book to read it, the act of stealing is itself a crime, regardless of what you do with the knowledge afterward.

Judge Alsup was going to let a jury decide the "pirated acquisition" question, and Anthropic rushed to settle — because no one knows how much a jury might award in damages. $1.5 billion is a staggering sum, but it beats a jury slapping you with $50 billion or $100 billion out of anger.

So the real meaning of this ruling is: the legality of AI training is confirmed, but the compliance cost of data acquisition just went up. From now on, LLM companies can't just scrape books from pirate sites willy-nilly. They either pay for licenses or use only legally available data sources.

Is $1.5 Billion Actually Expensive?

$1.5 billion sounds like an astronomical number. But do the math, and it might be "cheaper" than you think.

500,000 works, $3,000 per work. For a bestseller, $3,000 might not even cover a single year's royalties. But for the vast majority of ordinary writers, a book might never earn $3,000 in its entire lifetime. Many niche authors, academics, and poets — whose books might have sold only a few hundred copies total — are suddenly getting $3,000 in compensation. It's like money falling from the sky.

And for Anthropic? What's $1.5 billion? Its most recent funding round valued the company at over $60 billion; $1.5 billion is about 2.5% of that. Put another way: Anthropic just signed a $19 billion, 20-year compute deal with TeraWulf — the $1.5 billion settlement is less than one-tenth of a single compute contract.

More importantly, this settlement didn't set a precedent. Because it's a settlement, not an appeals court ruling, it's not binding on other cases. Google, Meta, and OpenAI are still fighting their own copyright battles, and every judge can reach their own conclusions. Just last week, major publishers including Hachette, Cengage, and Elsevier, along with author Scott Turow, filed a class-action lawsuit against Google, alleging the company used their copyrighted works to train Gemini.

"Training is fair use, but pirating the data is illegal — the essence of this ruling is: it doesn't ban AI, it just asks AI companies to pay for things like normal people."— The Dawn Vision Editorial Desk

So is $1.5 billion expensive or cheap? It depends on your perspective. From the writers' angle, $3,000 per book is certainly better than getting nothing, but it's probably a drop in the bucket compared to the value AI companies extract from their content. From Anthropic's angle, $1.5 billion buys a vague confirmation that training is legal and avoids a potentially astronomical jury award — not a bad deal at all.

An Unexpected Win for Chinese Model Companies

After this ruling, the ones secretly smiling might not be Anthropic — they might be China's LLM companies.

Why? Because Chinese LLM companies have faced stricter data compliance requirements from day one. China's regulatory environment means domestic LLM companies have always been relatively "disciplined" about training data acquisition — you can't just scrape everything, you can't use whatever you want, and copyright issues have been a sword of Damocles from the start. The result? Chinese LLM companies might actually have lower data compliance costs than their American counterparts — because they followed the rules from the beginning and never accumulated massive legal risk.

Right as this ruling dropped, another story is worth noting: The Verge published a long piece headlined "China delivers a one-two punch to America's AI dominance." The article argues that Moonshot's Kimi K3 and Alibaba's Qwen are pincering American AI companies from two directions — one hitting the high end with a closed-source model approaching GPT-5 and Claude performance, the other hitting the ecosystem with a fully open-source model. Together, they're rapidly narrowing America's AI lead.

On prediction market Polymarket, the probability of Anthropic reaching a $1.5 trillion valuation by year's end has dropped sharply, with traders directly linking the shift to Kimi K3's release. Meanwhile, Hong Kong-listed Zhipu AI surged over 30% in afternoon trading — the company launched a 1GW domestic compute center and completed its acquisition of Zhongke Jiahe.

The deeper impact of the copyright ruling is this: if American AI companies have to pay higher compliance costs for training data while Chinese companies are already lean on data compliance, the cost gap between them narrows further. Everyone used to assume American AI companies were "faster," but the people running in front are also carrying more legal landmines. As those mines start going off, the people behind might just catch up.

And let's not forget: China has the world's largest Chinese-language corpus and a unique regulatory framework. China's generative AI management measures were introduced early, with relatively clear boundaries for data compliance. From this perspective, markets that regulate first might benefit first — like how places that build roads first get traffic first.

Endgame: The Era of Data Costs Has Arrived

Put the $1.5 billion copyright settlement, the rise of Chinese models, and the evolution of the MCP protocol together, and a clear inflection point emerges: AI competition is shifting from "who has the bigger model" to "who has cleaner data, more controllable costs, and a better ecosystem."

2023 was about parameter count. 2024 was about performance. 2025 was about price. What's 2026 about? Compliance, data quality, infrastructure.

This ruling will have several profound effects:

First, the price of high-quality copyrighted data will skyrocket. Previously, publishers' content was taken by AI companies for free. Now, with the $1.5 billion precedent, publishers will grow a spine — and licensing fees will definitely go up. The content controlled by the Big Four publishers (Penguin Random House, HarperCollins, Simon & Schuster, Hachette) will become the "oil drilling rights" of the AI era. Building a high-quality model will first require negotiating licenses with these publishers — which circles back to the old rule: he who has the money calls the shots.

Second, the relative advantage of open-source models will grow. Closed-source model companies need to pay massive costs for training data, but the open-source community naturally has more legal data sources — code voluntarily contributed by developers, open-access academic papers, Creative Commons-licensed content, and more. If closed-source costs keep rising due to copyright issues while open-source performance keeps approaching, what happens to the market landscape? The answer speaks for itself.

Third, synthetic data will become a serious field. If real-world data has copyright problems, what about training AI with AI-generated data? That's the synthetic data approach. There's already plenty of research exploring using LLMs to generate training data for next-generation models. If this path works, copyright won't be the bottleneck anymore — AI will grow its own food. Of course, synthetic data has its own problems, like model collapse and lack of diversity, but compared to billions in copyright fees, these technical problems are cheap.

Fourth, global AI regulation enters the "detail-sorting" phase. South Korea's AI Framework Act takes effect today (July 21). The EU AI Act went fully into effect on July 1. China's generative AI management measures have been in place for a while. The US is still figuring out boundaries through lawsuits. Different regions have different regulatory frameworks and different compliance costs — which itself becomes a form of competitiveness. Wherever regulation is clear, predictable, and reasonably priced, that's where the AI industry will cluster.

The $1.5 billion settlement isn't the end — it's just the beginning. It's the first landmark event marking AI industry's shift from "wild growth" to "compliant development." There will be more copyright lawsuits, more regulatory policies, more compliance costs. AI companies can't just stare at model performance and funding valuations anymore. The legitimacy of the data supply chain is becoming the next core competitive advantage.

After all, in an era where data is oil, you first need to make sure the oil you're using — isn't stolen.

See you tomorrow.

训练本身是合理使用,但盗版获取违法——这个判决的精髓在于:它没有禁止AI,它只是要求AI公司像正常人一样付钱买东西。

—— Dawn Vision编辑部

Training is fair use, but pirating the data is illegal — the essence of this ruling is: it doesn't ban AI, it just asks AI companies to pay for things like normal people.

— The Dawn Vision Editorial Desk
Anthropic · 15亿美元 · 版权和解 · fair use · 合理使用 · AI训练数据 · 版权诉讼 · 中国大模型 · 数据合规 · 合成数据
Anthropic · $1.5 billion · copyright settlement · fair use · AI training data · copyright lawsuit · Chinese LLMs · data compliance · synthetic data
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 22 个源信号生成,经编辑部人工审核。素材来源:TechCrunch、Reuters、The Verge、Ars Technica、36氪、Polymarket。

This article was generated by the Dawn Vision cognitive engine processing 22 source signals, with human editorial review. Sources: TechCrunch, Reuters, The Verge, Ars Technica, 36Kr, Polymarket.