100毫秒。这是人类感知到"对方反应迟钝"的临界值。
7月6日,阿里千问大模型正式升级实时语音识别模型Fun-ASR-Realtime——一款首字延迟控制在百毫秒级别、识别准确率接近离线模型水平的流式语音识别大模型。更重要的是,单模型支持16种方言和30种语言,离线模式模型Fun-ASR-Flash也同步上线。API已接入阿里云百炼平台,开发者可以直接调用。
为什么实时语音这么重要?
语音是人类最自然的交互方式,但AI语音交互一直有一个致命问题:延迟。你跟Siri说话,它要等你说完、过一两秒才响应,那种等待感让人瞬间出戏。你跟真人对话的时候,对方是边听边想边回应的,延迟通常在200毫秒以内。AI如果做不到这个水平,就永远只是一个"语音命令识别器",而不是真正的对话伙伴。
百毫秒是什么概念?比人类眨眼的速度(约300毫秒)还快。这意味着你话说到一半,AI已经开始处理了;你刚说完最后一个字,AI几乎同时给出回应。这种"无缝衔接"的对话体验,是AI Agent、智能硬件、车载系统、客服场景的刚需。
实时语音的技术难度在于它是一个"不可能三角":延迟、准确率、模型大小三者很难同时满足。传统离线模型准确率高但延迟大,传统流式模型延迟低但准确率差尤其是在方言、噪声、口音场景下。Fun-ASR-Realtime的卖点是在百毫秒延迟的前提下,把准确率做到了接近离线模型的水平。
语音成为大模型竞争新焦点
2026年下半年,实时语音正在成为大模型厂商新的竞争焦点。
OpenAI的GPT-4o在2024年就推出了实时语音对话功能,但延迟和准确率一直在优化;Google Gemini也在实时多模态上加码;苹果iOS 27 beta本周开放了Siri语速和表现力自定义——虽然只是小步迭代,但说明苹果也在让AI助手说话更自然。国内方面,微信"小微"、豆包、Kimi等都在强化语音交互能力。
千问这次选择的切入点很务实:方言和多语言。中国是一个方言大国,粤语、四川话、上海话、河南话、湖南话……很多中老年人在说普通话时带有浓重口音,通用ASR模型在这些场景下准确率经常断崖式下跌。Fun-ASR-Realtime单模型支持16种方言,这对于下沉市场、老年用户、地方场景(如方言客服、车载方言交互)有实际价值。30种语言的覆盖也为出海场景做了准备。
从产业角度看,实时语音大模型的竞争是AI入口战争的延续。2023年大家在卷文本对话,2024-2025年卷图像和视频,2026年下半年开始卷实时语音和多模态交互。因为当AI Agent真正进入家庭、进入汽车、进入工作场景时,语音是最高频、最自然的交互方式。一个语音反应迟钝、听不懂方言的AI助手,就像一个耳朵不好、反应迟缓的同事——没人愿意用。
离线版本Fun-ASR-Flash的同步上线也值得关注。实时模型需要云端算力支持,但离线模型可以跑在设备本地——这意味着智能手表、车载终端、IoT设备等没有稳定网络或对隐私敏感的场景,也能用上高质量的语音识别。端云结合会是语音AI的主流部署模式。
"AI语音交互延迟每降低100毫秒,用户的使用意愿可能提升一倍——因为对话的'自然感'是一个非线性阈值。"—— 一位语音AI创业者的观察
大模型的竞争正在从"谁更聪明"(benchmark分数)转向"谁更好用"(交互体验)。百毫秒延迟的语音交互,可能是2026年下半年AI产品体验的分水岭。当AI说话不再像机器人、反应不再像掉线,真正的对话式AI时代才算到来。
明天见。
100 milliseconds. That's the threshold at which humans perceive "the other person is slow to respond."
On July 6, Alibaba's Qwen LLM officially upgraded its real-time speech recognition model Fun-ASR-Realtime — a streaming speech recognition LLM with first-token latency controlled at the 100ms level and recognition accuracy approaching offline model performance. More importantly, the single model supports 16 dialects and 30 languages. An offline model, Fun-ASR-Flash, also launched simultaneously. APIs are available on Alibaba Cloud's Bailian platform for developers to call directly.
Why Is Real-Time Speech So Important?
Speech is humanity's most natural interaction mode, but AI voice interaction has always had one fatal problem: latency. You talk to Siri, it waits until you finish, then responds after a second or two — that wait instantly breaks immersion. When you talk to a real person, they listen, think, and respond simultaneously, typically under 200ms. If AI can't hit that bar, it'll never be more than a "voice command recognizer," not a true conversational partner.
What does 100ms feel like? Faster than the human blink (~300ms). It means the AI starts processing while you're still speaking; the moment you finish your last word, the AI responds almost simultaneously. That "seamless" conversational experience is a hard requirement for AI Agents, smart hardware, in-car systems, and customer service scenarios.
The technical challenge of real-time speech is an "impossible triangle": latency, accuracy, and model size are extremely difficult to optimize simultaneously. Traditional offline models have high accuracy but high latency; traditional streaming models have low latency but poor accuracy, especially with dialects, noise, and accents. Fun-ASR-Realtime's selling point is achieving near-offline accuracy while maintaining 100ms-level latency.
Speech Becomes the New LLM Battleground
In H2 2026, real-time speech is becoming the new competitive focus for LLM vendors.
OpenAI's GPT-4o launched real-time voice conversation back in 2024, but latency and accuracy have been under continuous optimization; Google Gemini is doubling down on real-time multimodal; Apple's iOS 27 beta opened Siri speaking rate and expressiveness customization this week — small iterations, but they signal Apple is also pushing AI assistants to sound more natural. Domestically, WeChat's "Xiaowei," Doubao, Kimi and others are all strengthening voice interaction capabilities.
Qwen's entry point is pragmatic: dialects and multilingual support. China is a dialect-rich country — Cantonese, Sichuanese, Shanghainese, Henan dialect, Hunan dialect... many middle-aged and elderly people speak Mandarin with heavy accents, and generic ASR models often see accuracy cliff-dive in these scenarios. Fun-ASR-Realtime's single-model support for 16 dialects has practical value for lower-tier markets, elderly users, and regional scenarios (like dialect customer service, in-car dialect interaction). Coverage of 30 languages also prepares the model for overseas expansion.
From an industry perspective, the real-time speech LLM competition is a continuation of the AI entry-point wars. In 2023 everyone competed on text dialogue; 2024–2025 was image and video; H2 2026 is turning to real-time speech and multimodal interaction. Because when AI Agents truly enter homes, cars, and workplaces, voice is the highest-frequency, most natural interaction mode. An AI assistant with sluggish voice responses that can't understand dialects is like a hard-of-hearing, slow-reacting coworker — nobody wants to use it.
The simultaneous launch of the offline version, Fun-ASR-Flash, is also noteworthy. Real-time models require cloud compute support, but offline models can run on-device — meaning smartwatches, in-car terminals, IoT devices and other scenarios without stable networks or with privacy sensitivity can also use high-quality speech recognition. Edge-cloud combination will be the dominant deployment pattern for voice AI.
"Every 100ms reduction in AI voice interaction latency could double user willingness to engage — because conversational 'naturalness' is a non-linear threshold."— A voice AI founder's observation
LLM competition is shifting from "who's smarter" (benchmark scores) to "who's more usable" (interaction experience). 100ms-latency voice interaction could be the watershed for AI product experience in H2 2026. When AI stops sounding like a robot and responding like a dropped connection, the true conversational AI era will have arrived.
See you tomorrow.
AI语音交互延迟每降低100毫秒,用户的使用意愿可能提升一倍——因为对话的'自然感'是一个非线性阈值。
—— 一位语音AI创业者的观察
Every 100ms reduction in AI voice interaction latency could double user willingness to engage — because conversational 'naturalness' is a non-linear threshold.
— A voice AI founder's observation
Qwen · Fun-ASR-Realtime · real-time speech recognition · ASR · 100ms latency · dialect recognition · voice interaction · LLM competition · Alibaba
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 8 个源信号生成,经编辑部人工审核。素材来源:新浪财经、IT之家、搜狐、阿里云百炼。
This article was generated by the Dawn Vision cognitive engine processing 8 source signals, with human editorial review. Sources: Sina Finance, IT Home, Sohu, Alibaba Cloud Bailian.