具身智能 · 机器人

LingBot-World 2.0开源
世界模型首次小时级生成

LingBot-World 2.0 Open Sources
World Models Hit Hour-Long Generation

蚂蚁灵波开源LingBot-World 2.0和LingBot-Video,支持小时级连续生成、720p/60fps实时交互,世界模型从分钟级迈入小时级,具身智能有了"长期记忆"的世界底座。

Ant LingBot open sources LingBot-World 2.0 and LingBot-Video, supporting hour-long continuous generation at 720p/60fps real-time interaction — world models move from minutes to hours, giving embodied AI a world foundation with "long-term memory."

No.012 2026.07.10 约 5 分钟阅读 ~5 min read

AI生成的世界,终于不再"断片"了。

7月9日,蚂蚁集团旗下蚂蚁灵波科技正式开源新一代实时交互世界模型LingBot-World 2.0(又称LingBot-World-Infinity),同步开源全球首个具身智能专属MoE视频模型LingBot-Video。前者支持小时级连续稳定生成、720p/60fps高清实时交互,后者是专为具身智能训练设计的MoE架构视频生成模型——两项开源,把世界模型的天花板从分钟级抬到了小时级。

从"几分钟"到"几小时"意味着什么

理解这次升级的分量,需要回到世界模型本身的技术瓶颈。

世界模型(World Model)是具身智能的核心基础设施——机器人要在真实世界中行动,需要一个能预测物理世界演化的内部模型,就像人类大脑能预测"如果我推这个杯子,它会倒、水会洒"一样。但此前开源的世界模型有一个致命问题:生成时间短,长了就崩。大多数模型只能稳定生成几秒到几分钟的一致场景,时间一长就会出现物体消失、物理规律错乱、场景跳变——也就是业内说的"断片"。

LingBot-World 2.0给出的数据很硬:小时级连续稳定生成,720p分辨率、60fps帧率,支持AI原生多人实时交互。这意味着AI生成的世界可以连续运转一个小时以上,保持物理一致性、物体恒常性和场景连贯性——就像一个真实的小世界,而不是一段循环播放的GIF。

为什么这件事重要?因为具身智能的训练和评测需要长时间跨度的交互环境。一个机器人在真实世界里执行一项任务可能需要几十分钟甚至几小时,如果世界模型只能撑几分钟,就无法支撑完整的任务训练和评估。小时级生成能力让世界模型第一次具备了作为"具身智能训练场"的实用价值——机器人可以在这个虚拟世界里接受长时间、多步骤的任务训练,而不用担心世界突然"崩塌"。

MoE视频模型补全具身数据短板

同日开源的LingBot-Video填补了另一个关键空白:具身智能训练数据

具身智能的训练需要大量第一人称视角的交互视频数据——机器人看到什么、做了什么、结果如何。但真实世界的具身数据采集成本极高——需要机器人在真实环境中操作数小时,收集传感器数据,而且每次环境变化都要重新采集。这和大语言模型可以轻松从互联网获取万亿token文本形成了鲜明对比。

LingBot-Video是全球首个具身智能专属的MoE(Mixture of Experts)视频生成模型,可以批量生成高质量的具身交互视频数据,用于训练和评测具身智能体。量子位的报道中提到,这是全球首个具身专属的MoE视频模型。用合成数据补充甚至替代真实采集数据,是具身智能规模化的关键一步——就像大模型用合成数据提升性能一样,具身智能也需要自己的"数据引擎"。

蚂蚁灵波在一周内连续开源两款重磅模型(上周开源了LingBot-Vision视觉模型),节奏密集得不像一家大厂的作风。这背后的战略意图很清晰:通过开源建立具身智能世界模型的事实标准,吸引研究者和开发者在LingBot生态上构建具身应用,就像HuggingFace在NLP领域的位置一样。

世界模型是具身智能的"操作系统"——谁定义了世界模型,谁就定义了机器人如何理解和交互这个世界。小时级生成能力的突破,让具身智能从"实验室玩具"向"实用系统"又迈进了一大步。

明天见。

AI-generated worlds have finally stopped "blacking out."

On July 9, Ant Group's LingBot technology officially open-sourced its next-generation real-time interactive world model LingBot-World 2.0 (aka LingBot-World-Infinity), alongside LingBot-Video, the world's first MoE video model purpose-built for embodied intelligence. The former supports hour-long continuous stable generation at 720p/60fps high-definition real-time interaction; the latter is an MoE-architecture video generation model designed specifically for embodied intelligence training. Together, these two open-source releases push the world model ceiling from minutes to hours.

What Going From "Minutes" to "Hours" Actually Means

To understand the weight of this upgrade, return to the fundamental technical bottleneck of world models.

World models are the core infrastructure of embodied AI — for a robot to act in the real world, it needs an internal model that can predict how the physical world evolves, much like the human brain predicts "if I push this cup, it'll tip over and water will spill." But prior open-source world models had one fatal flaw: short generation windows before collapse. Most models could only stably generate a few seconds to minutes of consistent scenes; beyond that, objects vanished, physics broke, and scenes jumped — what the field calls "blacking out."

LingBot-World 2.0's numbers are concrete: hour-long continuous stable generation, at 720p resolution and 60fps, supporting AI-native multi-person real-time interaction. That means AI-generated worlds can run continuously for over an hour while maintaining physical consistency, object permanence, and scene coherence — like a real miniature world, not a looping GIF.

Why does this matter? Because embodied AI training and evaluation require long-horizon interactive environments. A robot performing a real-world task might take tens of minutes or hours. If a world model can only last minutes, it can't support complete task training and evaluation. Hour-long generation capability makes world models viable as "embodied AI training grounds" for the first time — robots can undergo long, multi-step task training in these virtual worlds without the world suddenly "collapsing."

MoE Video Model Fills the Embodied Data Gap

LingBot-Video, open-sourced the same day, fills another critical gap: embodied AI training data.

Embodied AI training requires massive amounts of first-person interactive video data — what the robot sees, what it does, and what happens as a result. But real-world embodied data collection is extremely expensive — requiring robots to operate for hours in real environments, collecting sensor data, with re-collection needed every time the environment changes. This stands in stark contrast to large language models, which can easily pull trillions of text tokens from the internet.

LingBot-Video is the world's first embodied-intelligence-specific MoE (Mixture of Experts) video generation model, capable of batch-generating high-quality embodied interaction video data for training and evaluating embodied agents. As QbitAI noted, this is the first MoE video model dedicated to embodied intelligence. Using synthetic data to supplement or even replace real collected data is a critical step for embodied AI to scale — just as large models use synthetic data to improve performance, embodied AI needs its own "data engine."

Ant LingBot has open-sourced two major models within a week (LingBot-Vision last week), a cadence uncharacteristically aggressive for a big company. The strategic intent is clear: establish a de facto standard for embodied world models through open source, attracting researchers and developers to build embodied applications on the LingBot ecosystem — claiming a position similar to HuggingFace in NLP.

World models are the "operating system" of embodied AI — whoever defines the world model defines how robots understand and interact with the world. The hour-long generation breakthrough pushes embodied AI another major step from "lab toy" toward "practical system."

See you tomorrow.

蚂蚁灵波 · LingBot-World 2.0 · LingBot-World-Infinity · 小时级生成 · 720p/60fps · 世界模型 · 具身智能 · LingBot-Video · MoE视频模型 · 开源 · 具身训练数据
Ant LingBot · LingBot-World 2.0 · LingBot-World-Infinity · hour-long generation · 720p/60fps · world model · embodied AI · LingBot-Video · MoE video model · open source · embodied training data
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 13 个源信号生成,素材来源:IT之家、量子位、今日头条。

This article was generated from 13 source signals processed by the Dawn Vision cognitive engine. Sources: IT Home, QbitAI, Toutiao.