AI音乐圈最害怕的那件事,终于发生了。
7月15日,一起黑客入侵事件把Suno推上了风口浪尖。黑客通过一名员工的凭证进入Suno的源代码系统,发现了一个惊人的事实:Suno的AI音乐模型,是用从YouTube、Genius、Deezer等平台爬取的数百万首歌曲训练的。
404 Media首先报道了这个消息。黑客泄露的内部文件显示,Suno的训练数据集里包含了数十年的音乐内容——从经典老歌到最新流行曲,从独立音乐到排行榜冠军,几乎无所不包。而这些音乐,Suno从来没有获得过版权方的授权。
这件事之所以重要,不只是因为Suno是AI音乐赛道的头部公司,更因为它是AI生成内容版权问题的又一次集中爆发。
为什么大家都在「偷偷爬」?
说句实话,Suno不是第一个这么干的,也不会是最后一个。
AI音乐、AI绘画、AI写作、AI视频——几乎所有生成式AI模型,训练数据都是「爬」来的。区别只在于:有的公司被发现了,有的没被发现;有的公司被起诉了,有的没被起诉。
为什么大家都要偷偷爬?因为正规授权太贵了。如果你要训练一个AI音乐模型,想从三大唱片公司(环球、索尼、华纳)正规拿授权,那授权费可能是几十亿甚至上百亿美元——这还没算上和无数独立音乐人谈授权的时间成本。大多数创业公司根本拿不出这笔钱。
那为什么版权方不主动授权给AI公司?因为他们还没想好怎么定价。授权便宜了,担心AI把自己的饭碗砸了;授权贵了,AI公司买不起。两边就在这种博弈中僵持着——AI公司偷偷爬,版权方睁一只眼闭一只眼,等着哪天攒够了筹码再秋后算账。
Suno这次被黑,相当于把这层窗户纸捅破了。以前大家都是「公开的秘密」,现在变成了「公开的证据」。版权方要是想告,现在有现成的证据了。
AI内容的版权困局,有解吗?
这件事的结局,大概率是Suno赔钱和解。但更深层的问题还在:AI生成内容的版权,到底应该怎么算?
现在的局面是一笔糊涂账。美国这边,法院判了AI生成的图片不受版权保护(因为没有人类作者),但没说训练数据合不合法;欧盟的AI法案说了一些原则,但细则还没出来;中国的《生成式人工智能服务管理暂行办法》要求训练数据合法,但怎么算「合法」也没说清楚。
可能的出路有几条:
第一条路:强制授权集体管理。就像KTV版权一样——AI公司不用一家一家去谈,直接给版权集体管理组织交钱,获得一揽子授权。简单粗暴,但有效。
第二条路:公平使用(Fair Use)扩大解释。美国法院可能会把AI训练认定为「公平使用」——就像搜索引擎索引网页一样,是合理的。这条路对AI公司最有利,但对创作者最不利。
第三条路:分成模式。AI公司按收入比例分给创作者,用得越多分的越多。听起来很美,但怎么统计「哪首歌对模型贡献了多少」?这是个技术难题。
不管走哪条路,有一点是肯定的:AI生成内容的「野蛮生长」时代,快要结束了。Suno这次事件只是一个开始,接下来会有更多的诉讼、更多的法规、更多的博弈。AI内容创作者们,做好准备吧。
明天见。
The thing the AI music industry feared most has finally happened.
On July 15, a hacking incident pushed Suno into the spotlight. A hacker gained access to Suno's source code system through an employee's credentials and discovered a shocking truth: Suno's AI music model was trained on millions of songs scraped from platforms like YouTube, Genius, and Deezer.
404 Media first broke the story. Internal files leaked by the hacker show Suno's training dataset contains decades of music — from classic oldies to the latest pop hits, from indie tracks to chart-toppers, covering almost everything. And Suno never obtained authorization from copyright holders for any of it.
The reason this matters isn't just that Suno is a leader in AI music — it's that this is yet another concentrated eruption of the copyright problem in AI-generated content.
Why Is Everyone "Scraping in Secret"?
Let's be honest: Suno isn't the first to do this, and it won't be the last.
AI music, AI art, AI writing, AI video — nearly all generative AI models are trained on scraped data. The only difference is: some companies got caught, some didn't; some got sued, some didn't.
Why does everyone scrape in secret? Because proper licensing is too expensive. If you want to train an AI music model and properly license from the three major labels (Universal, Sony, Warner), the licensing fees could be billions of dollars — and that's before accounting for the time cost of negotiating with countless independent artists. Most startups simply can't afford it.
Why don't rights holders proactively license to AI companies? Because they haven't figured out pricing yet. License too cheap and they worry AI will cannibalize their own business. License too expensive and AI companies can't afford it. Both sides are locked in a stalemate — AI companies scrape in secret, rights holders look the other way, waiting for the right moment to settle scores later.
Suno getting hacked is like poking a hole through this paper wall. Before, it was an "open secret." Now it's "public evidence." If rights holders want to sue, they've got the evidence ready-made.
Is There a Solution to AI Content's Copyright Dilemma?
The outcome of this will probably be Suno paying damages and settling. But the deeper question remains: how should copyright for AI-generated content actually work?
Right now the situation is a mess. In the U.S., courts have ruled that AI-generated images aren't copyrightable (no human author), but haven't said whether training data is legal. The EU AI Act lays out some principles but the details aren't out yet. China's "Generative AI Services Administration Measures" require legal training data, but what counts as "legal" isn't clearly defined either.
There are a few possible paths forward:
Path one: compulsory collective licensing. Like KTV copyright — AI companies don't negotiate one by one, they just pay a collective rights management organization and get blanket licensing. Crude but effective.
Path two: expanded Fair Use interpretation. U.S. courts could rule that AI training counts as "fair use" — just like search engines indexing web pages. This path is best for AI companies but worst for creators.
Path three: revenue sharing model. AI companies share revenue proportionally with creators — the more something is used, the more it earns. Sounds great, but how do you track "how much each song contributed to the model"? That's a technical nightmare.
Whichever path we take, one thing is certain: the "wild growth" era of AI-generated content is coming to an end. The Suno incident is just the beginning. More lawsuits, more regulations, more battles are coming. AI content creators, get ready.
See you tomorrow.
AI生成内容的「野蛮生长」时代,快要结束了。Suno这次事件只是一个开始。
—— Dawn Vision编辑部
The "wild growth" era of AI-generated content is coming to an end. The Suno incident is just the beginning.
— The Dawn Vision Editorial Desk
Suno · AI Music · Copyright Dispute · Training Data · YouTube · Generative AI · Content Creation