AI 设计 · 创意工具

Google DeepMind发布Lyria 3.5
音乐性/歌词/人声全面升级,AI音乐进"发片级"

Google DeepMind Launches Lyria 3.5
Musicality/Lyrics/Vocals Upgraded, Release-Ready

Google DeepMind 7月29日发布Lyria 3.5音乐生成模型,在Flow Music上线,音乐性、歌词质量、人声情感、创意控制大幅提升。

Google DeepMind released Lyria 3.5 music generation model on July 29, live in Google Flow Music, with major improvements across musicality, lyrics, vocal emotion, and creative control.

No.025 2026.07.30 约 5 分钟阅读 ~5 min read

AI生成的图片已经让摄影师焦虑,AI生成的视频已经让影视公司焦虑,现在轮到音乐行业了。

7月29日,Google DeepMind发布Lyria 3.5音乐生成模型,同步在Google Flow Music上线。这次升级不是小修小补——官方在四个维度做了系统性提升:音乐性、歌词、人声、创意控制。看完官方演示,一个感受非常清晰:AI音乐的质量,已经从"Demo级"进入了"发片级"。

四个维度的升级,每个都打在痛点上

用过AI音乐工具的人都知道几个常见痛点:旋律像MIDI模板一样机械、歌词前言不搭后语、人声像机器人没有感情、生成后想调速度调时长只能重新生成。Lyria 3.5直接对着这四个痛点开刀。

音乐性提升:旋律结构更丰富、更复杂,听起来更自然。之前的AI音乐你听10秒就能听出"这是AI写的",因为和弦进行和旋律走向太"教科书"、太"正确",缺少真人创作时的"意外感"。Lyria 3.5在这个维度做了明显改进。

歌词质量增强:更高质量的歌词生成,prompt遵从度更好,结构意识更强。之前AI写歌词经常为了押韵硬凑词、主题跑偏、主歌副歌结构混乱。新版本可以生成结构完整、主题连贯、甚至有点文学性的歌词。

人声改进:更真实、更有情感细节的人声,发音也更准确。AI人声的" uncanny valley"(恐怖谷)效应一直是最难突破的——它唱准了所有音,但你就是觉得"没人味"。Lyria 3.5在情感表达和咬字上做了提升。

创意控制:可以更容易控制输出的速度(tempo)和时长。这听起来是小功能,但对实际创作极其重要——做短视频BGM需要15秒、做播客片头需要30秒、做完整歌曲需要3分钟,以前你只能反复生成碰运气,现在可以直接指定。

"AI生成图片已经骗过你的眼睛,AI生成视频已经骗过你的眼睛,现在AI生成音乐开始骗过你的耳朵——创作者的工具库又多了一件核武器。"—— Dawn Vision编辑部

AI音乐的"ChatGPT时刻"还有多远

Lyria 3.5当然不是终点。和专业音乐人相比,AI音乐在情感深度、即兴发挥、个人风格上还有不小差距。但拐点已经清晰可见。

想想AI图片的演化路径:2022年Midjourney v3刚出来时大家觉得"哇好厉害"但不会真的用在商业项目里;2023年Midjourney v5/v6出来后,广告、电商、自媒体开始大规模使用;2026年的今天,AI图片已经是内容创作的标配工具。AI音乐正在走同样的路——只是晚了大约18-24个月。

对独立创作者、短视频博主、播客制作人、游戏开发者来说,这是巨大的生产力释放:以前你找音乐制作人做一首BGM可能要几千块、等一周;现在你可以用Lyria 3.5在几分钟内生成几个版本,挑最合适的用。门槛被打下来了。

当然,版权问题、收入分配、真人音乐家的生存空间——这些问题不会因为技术进步自动解决。但技术的车轮不会等人。

明天见。

Sources · 参考来源

声明:本文为 Dawn Vision 基于公开信息的二次创作与独立分析,标题、观点、行文均为原创,仅供参考,不构成任何投资建议或决策依据。如有侵权请联系删除。

本文基于 Dawn Vision 认知引擎处理的 6 个源信号生成,经编辑部人工审核。素材来源:Google DeepMind官方博客。

相关入库笔记:Google DeepMind · Lyria 3.5 · AI音乐 · 音乐生成 · Flow Music · AI创意工具

AI-generated images already made photographers anxious; AI-generated video already made film companies anxious — now it's the music industry's turn.

On July 29, Google DeepMind released Lyria 3.5, live immediately in Google Flow Music. This wasn't a minor patch — the official announcement highlights systematic improvements across four dimensions: musicality, lyrics, vocals, and creative control. After watching the official demos, one feeling is clear: AI music quality has crossed from "demo-grade" to "release-grade."

Four Upgrades, Each Hitting a Real Pain Point

Anyone who's used AI music tools knows the common pain points: melodies as mechanical as MIDI templates, lyrics that don't cohere, vocals that sound robotic and emotionless, and wanting to adjust tempo/length meaning regenerating from scratch. Lyria 3.5 goes after all four.

Improved musicality: richer, more complex melodic structures that sound more natural. Earlier AI music you could spot within 10 seconds because chord progressions and melodic arcs were too "textbook," too "correct," lacking the "happy accidents" of human composition. Lyria 3.5 makes visible improvement here.

Enhanced lyrics: higher quality lyric generation with better prompt adherence and structural awareness. Previous AI lyrics often forced rhymes at the cost of meaning, wandered off-topic, or had messy verse-chorus structures. The new version produces lyrics with complete structure, coherent themes, even occasional literary quality.

Improved vocals: more realistic, emotionally nuanced vocals with better pronunciation. The uncanny valley in AI vocals was always the hardest barrier — it hits all the notes but you can tell "there's no human there." Lyria 3.5 improves emotional expression and diction.

Creative control: easier control over tempo and duration. Sounds like a small feature, but it's huge for real creation — short-video BGM needs 15 seconds, podcast intros need 30 seconds, full songs need 3 minutes; before you had to regenerate and hope; now you can specify directly.

"AI images already fooled your eyes, AI video already fooled your eyes — now AI music starts fooling your ears, and it writes lyrics, sings, controls tempo. Creators just added another nuke to their toolkit."—— The Dawn Vision Editorial Desk

How Far Is AI Music's 'ChatGPT Moment'?

Lyria 3.5 isn't the endpoint, of course. Compared to professional human musicians, AI music still has gaps in emotional depth, improvisation, and personal style. But the inflection point is visible.

Consider AI images' trajectory: when Midjourney v3 came out in 2022, people thought "wow cool" but wouldn't use it commercially; by Midjourney v5/v6 in 2023, ads, e-commerce, and social media adopted it at scale; in 2026, AI images are a standard content creation tool. AI music is walking the same path — just roughly 18-24 months behind.

For indie creators, short-video creators, podcast producers, game developers, this is a massive productivity unlock: commissioning a BGM from a human producer might cost thousands and take a week; now you can generate multiple versions in minutes with Lyria 3.5 and pick the best fit. Barriers are collapsing.

Copyright issues, revenue sharing, human musicians' livelihoods — these problems won't automatically solve themselves with technical progress. But the technological wheel doesn't wait.

See you tomorrow.

Sources · 参考来源

声明:本文为 Dawn Vision 基于公开信息的二次创作与独立分析,标题、观点、行文均为原创,仅供参考,不构成任何投资建议或决策依据。如有侵权请联系删除。

This article was generated by the Dawn Vision cognitive engine processing 6 source signals, with human editorial review. Sources: Google DeepMind official blog.

相关入库笔记:Google DeepMind · Lyria 3.5 · AI music · music generation · Flow Music · generative audio · AI creative tools

AI生成图片已经骗过你的眼睛,AI生成视频已经骗过你的眼睛,现在AI生成音乐开始骗过你的耳朵——而且它还能写歌词、唱歌、控制节奏。创作者的工具库又多了一件核武器。

—— Dawn Vision编辑部

AI images already fooled your eyes, AI video already fooled your eyes — now AI music starts fooling your ears, and it writes lyrics, sings, controls tempo. Creators just added another nuke to their toolkit.

—— The Dawn Vision Editorial Desk
Google DeepMind · Lyria 3.5 · AI音乐 · 音乐生成 · Flow Music · 生成式音频 · AI创意工具
Google DeepMind · Lyria 3.5 · AI music · music generation · Flow Music · generative audio · AI creative tools
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 6 个源信号生成,经编辑部人工审核。素材来源:Google DeepMind官方博客。

This article was generated by the Dawn Vision cognitive engine processing 6 source signals, with human editorial review. Sources: Google DeepMind official blog.