具身智能 · 机器人

智象发布具身世界模型
扰动适应0.692分登顶RoboColiseum

HiDream Launches Embodied World Model
Tops RoboColiseum Robustness at 0.692

智象发布具身世界模型HiDream-O1-Embodied,扰动适应指标0.692分登顶RoboColiseum。结合一个月前的交互式世界模型,原生全模态技术闭环基本跑通。

HiDream.ai launches embodied world model HiDream-O1-Embodied, topping RoboColiseum Robustness at 0.692. Combined with last month's interactive world model, the native full-modal loop is essentially complete.

No.048 2026.09.08 约 5 分钟阅读 ~5 min read

从模拟世界到真实世界的关键一步

智象未来(HiDream.ai)正式发布具身世界模型HiDream-O1-Embodied,主打机器人物理感知与动态预判能力。这不是一个孤立的模型发布——不到一个月前,智象刚推出交互式世界模型HiDream-O1-World并登顶WBench Navi分榜。一个负责「理解与推演」,一个负责「操作与执行」,两条线合起来,智象的原生全模态世界模型技术闭环基本跑通了:图像、视频、3D、动作模态打通,理解—推演—执行全流程贯通。

发布即登顶。HiDream-O1-Embodied首次参评具身智能评测平台RoboColiseum,就在最具挑战性的扰动适应(Robustness)子榜单上以0.692分拿下第一。这个榜单检验的不是温室里的理想表现,而是背景、光照、材质、视角、指令改写统统变化时,模型还能不能稳稳干活。

「下一代大模型竞争的关键,不在于单一模态能力的增长,而是从单模态走向多模态,并走向原生统一的全模态。」——梅涛,智象未来创始人兼CEO

三项硬核能力:从「一处失灵」到「局部受限仍可运行」

传统机器人的三大痛点,HiDream-O1-Embodied挨个破局。语言理解上,它突破了关键词匹配的局限——「把杯子拿过来」「给我拿个杯子」「杯子递我一下」,不管怎么说都能直达意图。视觉感知上,它不再依赖单一路视角,而是多视角协同互补,部分视觉信息出问题时,靠其他视角照样能理解场景、继续执行,从「一处失灵、全局失效」变成「局部受限、整体仍可运行」。

容错机制是最底层的支撑。多数模型训练用的是完美数据——清晰画面、标准视角、完整信息。真实世界哪有这么好的事?光照变了、镜头脏了、视线被挡了、信号不稳了,都是日常。HiDream-O1-Embodied在训练阶段就主动引入各种非理想条件,让模型在不完整、不稳定的信息里练出可靠判断力。

模型+数据:自己出题自己答的增长飞轮

为什么能扛扰动?根子在数据策略。智象搞了一套「真实基座+生成增强」的数据生产范式——高精度真人动作捕捉当底座,模型自己出手,在严格遵守物理约束的前提下切换背景、光照、物体形态,百倍级扩增数据。

这就形成了一个飞轮:模型既是考生也是出题人,根据自身弱点针对性生成训练样本,数据和模型互相驱动往上走。从HiDream-O1-Image到HiDream-O1-World再到HiDream-O1-Embodied,视觉模型、交互式世界模型、具身世界模型的家族矩阵已经铺开,下一步就是看这套原生全模态的方法论能在真实机器人身上走多远了。

明天见。

A Key Step from Simulated Worlds to the Real World

HiDream.ai has officially launched HiDream-O1-Embodied, an embodied world model focused on robot physical perception and dynamic prediction. This is not an isolated model release — less than a month ago, the company unveiled its interactive world model HiDream-O1-World and topped the WBench Navi leaderboard. One handles "understanding and reasoning," the other "manipulation and execution." Together, they essentially complete HiDream's native full-modal world model loop: image, video, 3D, and motion modalities are unified; the understand–reason–execute pipeline is end-to-end.

It debuted at the top. Entering the RoboColiseum embodied AI benchmark for the first time, HiDream-O1-Embodied took first place on the most challenging Robustness sub-leaderboard with a score of 0.692. This list measures not ideal performance in a lab greenhouse, but whether the model can reliably get work done when backgrounds, lighting, materials, camera angles, and instruction phrasings all change.

"The key to the next generation of large model competition is not the growth of single-modality capability, but the move from single to multi-modality, and toward natively unified full-modality." — Tao Mei, Founder and CEO of HiDream.ai

Three Core Capabilities: From "One Failure Breaks Everything" to "Partial Limits, System Still Runs"

HiDream-O1-Embodied addresses three classic pain points of traditional robots head-on. In language understanding, it goes beyond keyword matching — "bring me the cup," "pass a cup over," "hand me the cup" — however you phrase it, it grasps the intent. In visual perception, instead of relying on a single fixed viewpoint, it fuses multiple perspectives so that when part of the visual signal is compromised, other views still carry scene understanding and execution forward, shifting the paradigm from "one failure breaks everything" to "partial limits, system still runs."

Fault tolerance is the foundation. Most models train on perfect data — clean images, standard viewpoints, complete information. The real world is never that nice. Lighting shifts, lenses get smudged, sightlines are blocked, signals fluctuate — all everyday occurrences. HiDream-O1-Embodied actively introduces all kinds of non-ideal conditions during training, forcing the model to develop reliable judgment from incomplete, unstable information.

Model + Data: A Flywheel Where the Model Writes Its Own Exam

Why the robustness? It traces back to data strategy. HiDream has built a "real-base + generative-augmentation" data production paradigm — high-precision human motion capture forms the foundation, and the model itself generates hundred-fold data expansion by switching backgrounds, lighting, and object shapes while strictly preserving physical constraints.

This creates a flywheel: the model is both student and exam-writer, generating targeted training samples based on its own weaknesses, with data and model driving each other upward. From HiDream-O1-Image to HiDream-O1-World to HiDream-O1-Embodied, the family matrix of vision models, interactive world models, and embodied world models is now in place. The next question is how far this native full-modal methodology can go on real physical robots.

See you tomorrow.

下一代大模型竞争的关键,不在于单一模态能力的增长,而是从单模态走向多模态,并走向原生统一的全模态。

—— 梅涛,智象未来创始人兼CEO

The key to the next generation of large model competition is not the growth of single-modality capability, but the move from single to multi-modality, and toward natively unified full-modality.

— Tao Mei, Founder and CEO, HiDream.ai
具身世界模型·全模态闭环·扰动适应·RoboColiseum·生成增强数据·智象未来·梅涛
embodied world model · full-modal loop · robustness · RoboColiseum · generative data augmentation · HiDream.ai · Tao Mei
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 4 个源信号生成,经编辑部人工审核。素材来源:量子位。