9月15日,Google发布了Gemini 3.8 Live。这不是一次渐进式的模型升级,而是对话AI技术栈的一次重新定义——它第一次让AI做到了"边思考边说话"。
97种语言与语音质量登顶
Gemini 3.8 Live支持97种语言的自动切换。用户在对话中切换语言时,模型无需任何提示或设置,自动识别并无缝响应。这对全球化场景下的AI应用——客服、翻译、教育——意味着一个质的飞跃:一个模型覆盖近百种语言,不再需要为每种语言单独部署。
在语音质量评估中,Gemini 3.8 Live的表现同样亮眼。Speech Quality Index(SQI)82.6分,位列所有已发布模型第一。Big Bench Audio得分97.7%,τ-Voice测试68.6%。这些指标意味着,在"听起来是否像人"这个维度上,Gemini 3.8 Live已经逼近了人类对话的自然度。
边思考边说话:对话AI的范式突破
Gemini 3.8 Live最核心的技术突破是"边思考边语音"(Extended Thinking + Voice)能力。传统对话AI的流程是"思考→生成文本→转语音→播放",用户需要等待模型完成整个链条。Gemini 3.8 Live打破了这个顺序:模型在进行复杂推理的同时,已经开始输出语音——用户可以实时听到AI的"思考过程"。
与此同时,Gemini 3.8 Live支持后台执行工具调用。当模型在对话中需要查询数据库、调用API或执行代码时,这些操作在后台异步进行,不阻塞语音输出。用户感受到的是一个流畅的、像人一样"边想边聊"的交互体验。
SynthID水印与可追溯性
Google同步部署了SynthID水印技术。Gemini 3.8 Live生成的所有语音内容都携带不可见的数字水印,确保AI生成内容的可追溯性。在AI生成内容日益泛滥的当下,这项技术的意义超越了产品本身——它为AI内容的"来源认证"提供了一个可执行的技术方案。
Gemini 3.8 Live不是在追赶GPT或Claude——它在定义一个新赛道。当AI可以边想边说、97种语言无缝切换、后台同时执行复杂任务时,"对话AI"这个词已经不足以描述它了。Google正在把Gemini打造成一个多模态、多语言、多任务的实时AI操作系统。
明天见。
On September 15, Google released Gemini 3.8 Live. This isn't an incremental model upgrade — it's a redefinition of conversational AI's tech stack. For the first time, an AI can "think while speaking."
97 Languages, Voice Quality Crown
Gemini 3.8 Live supports automatic switching across 97 languages. When users switch languages mid-conversation, the model detects and responds seamlessly — no prompts, no settings. For global AI applications — customer service, translation, education — this is a qualitative leap: one model covering nearly a hundred languages, eliminating per-language deployments.
On voice quality metrics, Gemini 3.8 Live is equally dominant. Speech Quality Index (SQI) of 82.6 — first among all released models. Big Bench Audio score: 97.7%. τ-Voice test: 68.6%. These numbers mean that on the axis of "does it sound human," Gemini 3.8 Live is approaching the naturalness of human conversation.
Think While Speaking: A Paradigm Break
Gemini 3.8 Live's core technical breakthrough is "Extended Thinking + Voice". Traditional conversational AI follows a chain: think → generate text → convert to speech → play. Users wait for the entire pipeline. Gemini 3.8 Live breaks this sequence: the model begins voice output while still performing complex reasoning — users can hear the AI's "thinking process" in real time.
Simultaneously, Gemini 3.8 Live supports background tool invocation. When the model needs to query databases, call APIs, or execute code mid-conversation, these operations run asynchronously without blocking voice output. The user experiences a fluid, human-like "thinking out loud" interaction.
SynthID Watermarking and Traceability
Google has simultaneously deployed SynthID watermarking. All voice content generated by Gemini 3.8 Live carries invisible digital watermarks, ensuring traceability of AI-generated content. In an era of proliferating AI-generated media, this technology's significance transcends the product itself — it provides an executable technical solution for "source authentication" of AI content.
Gemini 3.8 Live isn't chasing GPT or Claude — it's defining a new race. When an AI can think while speaking, switch seamlessly across 97 languages, and execute complex tasks in the background simultaneously, "conversational AI" no longer describes it. Google is building Gemini into a multimodal, multilingual, multi-task real-time AI operating system.
See you tomorrow.
当AI可以边想边说、97种语言无缝切换、后台同时执行复杂任务时,"对话AI"这个词已经不足以描述它了。
—— Dawn Vision编辑部
When an AI can think while speaking, switch seamlessly across 97 languages, and execute complex tasks in the background, 'conversational AI' no longer describes it.
— The Dawn Vision Editorial Desk
Dawn Vision, Google, Gemini, 3.8 Live, voice AI, multilingual, SQI, SynthID, thinking-while-speaking, tool invocation
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 10 个源信号生成,经编辑部人工审核。素材来源:Google Blog、36氪、Unite.AI。
This article was generated by the Dawn Vision cognitive engine processing 10 source signals, with human editorial review. Sources: Google Blog, 36Kr, Unite.AI.