9月15日。
这是Cloudflare给全行业AI公司划出的最后期限。从这一天开始,Cloudflare的默认设置将阻断所有"混合用途"爬虫访问任何托管广告的页面——除非网站所有者手动调整设置,或者AI公司明确将搜索爬虫和AI训练/Agent用爬虫分离开来。
机器人流量已超过人类,内容生态难以为继
Cloudflare CEO Matthew Prince在公告中给出了一个关键数据:互联网上超过50%的流量已经是非人类流量,而且这个拐点比预期提前了一年到来。其中超过50%的AI爬虫流量花在重复抓取根本没有更新的页面上——既消耗出版商的带宽和服务器资源,又拿走了创作者的内容用来训练模型,却一分钱回报都没有。
Prince的表态很直接:"既然互联网上大多数流量已经是非人类的,我们必须更快行动,才能让一个可持续的生态系统出现。"
Cloudflare这次的新政主要有三层。第一是默认阻断混合爬虫:不明确区分搜索用途和AI训练/Agent用途的爬虫,9月15日后在Cloudflare免费版和新站点上默认被屏蔽。第二是Pay Per Use付费机制:从去年的Pay Per Crawl(按抓取次数付费)升级为按使用价值付费——不是AI爬虫抓一次就付一次钱,而是当AI公司用出版商的内容创造了价值时再付费。第三是合作伙伴试点:已经和Ceramic.ai、You.com达成合作,出版商可以选择加入,当内容出现在AI搜索结果或被AI助手引用时获得报酬。
Cloudflare还不点名地批评了"全球最大搜索引擎"(明显指Google):它能访问比其他AI公司多约2倍的信息,因为搜索巨头让网站所有者很难在保持搜索可发现性的同时不被用于AI训练。Google虽然提供了Google Extended让网站所有者可以选择退出AI训练,但核心的Googlebot同时为搜索和AI Overviews、AI Mode服务,属于Cloudflare定义的"混合用途爬虫"。
AI内容付费从"道德呼吁"进入"技术强制"阶段
过去两年,关于AI公司是否应该为训练数据付费的争论一直在持续,但基本停留在口水仗和诉讼层面。《纽约时报》起诉OpenAI、多家出版商联名上书、欧盟AI法案要求披露训练数据——这些动作要么进展缓慢,要么执行成本极高。
Cloudflare这次的不一样在于:它不是在立法层面呼吁,也不是在法院打官司,而是直接在网络基础设施层动了刀子。Cloudflare承载了全球相当比例的网站流量,它的默认设置改变,会直接影响数AI公司获取内容的能力。这是从"你应该付费"的道德呼吁,变成了"不付费你就拿不到"的技术强制。
这不是Cloudflare第一次在AI内容版权上出手。2024年它推出了对抗AI爬虫的工具,2025年推出了Pay Per Crawl交易市场。这次的新政是这条路线的延续:给网站所有者更多控制权,给AI公司明确的合规路径,给内容付费建立可操作的技术机制。
行业层面的连锁反应已经开始。同一周,美国参议员Elizabeth Warren和众议员Mary Gay Scanlon提出新法案,禁止AI公司出售用户的健康数据和位置信息给数据经纪商;音乐平台Tidal宣布从7月15日起不给100% AI生成的音乐支付版税,但不会 outright 禁止,而是打上标签。从网络基础设施到立法到内容平台,围绕AI数据和版权的规则体系正在快速成型。
9月15日这个deadline,会迫使AI公司做出选择:要么把搜索爬虫和训练/Agent爬虫分开、走付费合作路线;要么面临被大量网站默认屏蔽的风险。免费爬取全网内容训练模型的时代,可能真的要结束了。
September 15.
That's the deadline Cloudflare has drawn for AI companies across the industry. Starting that day, Cloudflare's default settings will block all "mixed-use" crawlers from accessing any ad-served pages -- unless website owners manually adjust settings, or AI companies clearly separate their search crawlers from AI training/Agent crawlers.
Bot Traffic Has Surpassed Humans, the Content Ecosystem Is Unsustainable
Cloudflare CEO Matthew Prince offered a key statistic in the announcement: over 50% of internet traffic is already non-human, and this inflection point arrived a year ahead of expectations. Over 50% of AI crawler traffic is spent re-scraping pages that haven't been updated at all -- draining publishers' bandwidth and server resources while taking creators' content to train models, with zero compensation in return.
Prince was direct: "Now that the majority of internet traffic is non-human, we must move faster for a sustainable ecosystem to emerge."
Cloudflare's new policy has three main layers. First is default blocking of mixed crawlers: crawlers that don't clearly separate search use from AI training/Agent use will be blocked by default on Cloudflare free plans and new sites after September 15. Second is the Pay Per Use mechanism: upgraded from last year's Pay Per Crawl (pay per crawl) to pay-by-value -- not paying every time an AI crawler hits a page, but paying when AI companies create value from publishers' content. Third is partner pilots: already partnered with Ceramic.ai and You.com, where publishers can opt in and get paid when their content appears in AI search results or is cited by AI assistants.
Cloudflare also criticized (without naming) the "world's largest search engine" (obviously Google): it can access roughly 2x more information than other AI companies because the search giant makes it hard for website owners to maintain search discoverability while opting out of AI training. While Google offers Google Extended for opting out of AI training, the core Googlebot serves both search and AI Overviews/AI Mode simultaneously, making it what Cloudflare defines as a "mixed-use crawler."
AI Content Payment Moves from "Moral Appeal" to "Technical Enforcement"
For the past two years, the debate over whether AI companies should pay for training data has raged on, but mostly stayed at the level of rhetoric and lawsuits. The New York Times suing OpenAI, multiple publishers signing open letters, the EU AI Act requiring training data disclosure -- these moves either progress slowly or carry enormous enforcement costs.
What's different this time is that Cloudflare isn't appealing at the legislative level or fighting in court; it's making the cut directly at the network infrastructure layer. Cloudflare carries a significant share of global website traffic, and changing its default settings will directly impact AI companies' ability to access content. This transforms "you should pay" from a moral appeal into "don't pay and you don't get it" technical enforcement.
This isn't Cloudflare's first move on AI content copyright. In 2024 it launched tools to counter AI crawlers; in 2025 it launched the Pay Per Crawl marketplace. This new policy is a continuation of that path: giving website owners more control, giving AI companies a clear compliance path, and establishing an actionable technical mechanism for content payment.
Industry chain reactions have already begun. The same week, US Senator Elizabeth Warren and Representative Mary Gay Scanlon introduced new legislation banning AI companies from selling users' health and location data to data brokers; music platform Tidal announced it would stop paying royalties on 100% AI-generated music starting July 15, though it won't outright ban it -- just label it. From network infrastructure to legislation to content platforms, the regulatory framework around AI data and copyright is rapidly taking shape.
The September 15 deadline will force AI companies to choose: either separate search crawlers from training/Agent crawlers and go the paid partnership route, or face the risk of being blocked by default across a vast number of websites. The era of freely scraping the entire web to train models may genuinely be coming to an end.
Cloudflare的新工具给了网站所有者更高的可见度和商业机会,也让爬虫意图清晰透明的AI公司受益。
—— Cloudflare CEO Matthew Prince
Cloudflare's new tools give website owners greater visibility and commercial opportunities, while also benefiting AI companies whose crawler intent is clear and transparent.
-- Matthew Prince, Cloudflare CEO
Cloudflare · AI Crawlers · Content Payment · Copyright · Pay Per Use · Bot Traffic · Training Data