你伸手拿起一个鸡蛋——你不用看,只用手指的力度变化就知道蛋壳有多脆、握到什么程度会碎。你拿起一杯水——指尖的温度和重量感告诉你水有多满、会不会洒。这些人类习以为常的“手感”,对机器人来说却是最难的课题之一。
7月,上海新智具身与复旦大学联合发布三份技术报告,通过构建包含超3万小时交互数据的开源触觉数据集,解决了触觉数据格式不统一的行业难题。这不是一个普通的数据集发布——它意味着具身智能的感知拼图,终于补上了“触觉”这关键的一块。
为什么触觉是具身智能的最尴尬短板
过去几年,具身智能的感知能力进步神速。视觉有了LingBot-Vision这样的空间原生基模,深度估计的精度一涨再涨;听觉有了各种语音识别和环境声分类模型。但触觉——这个人类感知世界最重要的通道之一——在机器人领域却一直是个“老大难”。
为什么触觉这么难?三个原因。
第一,触觉传感器太贵太脆弱。一个高精度的仿生触觉皮肤,成本可能几万甚至几十万元,而且一撞就坏。视觉传感器几百块钱一个,摔了也不心疼,触觉传感器却是“碰一下就几千块没了”。成本和耐用性的问题,导致触觉数据的采集规模一直上不去。
第二,触觉数据没有统一格式。视觉数据大家都用RGB图像,格式标准、工具丰富。但触觉呢?不同传感器厂商的数据格式完全不一样——有的输出压力分布矩阵,有的输出三维力向量,有的输出振动频率。各家数据不互通,研究成果没法复现,整个领域像是在各自为战。
第三,触觉和动作的耦合太紧密。视觉可以“站着看”,但触觉必须“动手摸”。你想采集触觉数据,就必须让机器人实际去接触物体——这意味着数据采集的成本是视觉的N倍,风险也是N倍。
3万小时的意义:不只是数据,是标准
新智具身和复旦这次发布的三份报告,最有价值的地方不是3万小时这个数字——而是他们建立了一套统一的触觉数据格式和采集规范。
这有点像ImageNet之于计算机视觉的意义。ImageNet最值钱的不是那1400万张图片,而是它建立了一个统一的基准——大家用同样的数据训练、同样的标准评测,研究才能在同一个轨道上加速进步。
3万小时触觉数据集的发布,可能会成为具身智能领域的“ImageNet时刻”。有了统一的数据格式和基准,不同团队的研究成果才能对比、才能复现、才能站在彼此的肩膀上。更重要的是,有了大规模真实触觉数据,机器人才能真正学会“怎么用力”——握鸡蛋用多大劲、拧螺丝用多大扭矩、端杯子怎么保持水平、揉面团怎么控制力度变化。
WAIC 2026上,具身智能与智算并列为两大核心赛道,超过200家企业、200余款具身智能终端、300多台真机集中亮相。但明眼人都看得出来:大部分机器人还停留在“能走能说”的阶段,真正“能干活”的没几个。干不了活的核心原因,不是脑子不够聪明(大模型已经够强了),也不是眼睛不够好用(视觉模型进步很快)——而是手不够灵巧、触感不够精细。
触觉这一关过不了,具身智能就forever confined to dancing on stage,进不了工厂和家庭。3万小时的数据集,是一个开始,但它标志着行业终于开始认真对待这个“最尴尬的短板”了。
明天见。
You reach out and pick up an egg — you don't need to look. The changing pressure in your fingers alone tells you how brittle the shell is, how hard you can squeeze before it breaks. You pick up a glass of water — the temperature and weight at your fingertips tell you how full it is, whether it might spill. These “senses of touch” that humans take for granted are one of the hardest problems in robotics.
In July, Shanghai Xinzhi Embodied and Fudan University jointly released three technical reports, building an open-source tactile dataset containing over 30,000 hours of interaction data, solving the industry-wide problem of fragmented tactile data formats. This isn't just another dataset release — it means the perception puzzle of embodied AI has finally gained the crucial piece that is “touch.”
Why Touch Is Embodied AI's Most Embarrassing Shortcoming
In the past few years, the perceptual capabilities of embodied AI have advanced by leaps and bounds. Vision has space-native foundation models like LingBot-Vision, with depth estimation accuracy rising and rising. Hearing has various speech recognition and environmental sound classification models. But touch — one of the most important channels through which humans perceive the world — has long been a persistent headache in robotics.
Why is touch so hard? Three reasons.
First, tactile sensors are too expensive and too fragile. A high-precision biomimetic tactile skin can cost tens or even hundreds of thousands of yuan, and it breaks on impact. Visual sensors cost a few hundred bucks each — you drop one, no big deal. But a tactile sensor? “Bump it once, and thousands of yuan are gone.” Cost and durability issues have kept tactile data collection scale from growing.
Second, there's no unified format for tactile data. For visual data, everyone uses RGB images — standard format, rich tooling. But touch? Data formats vary completely across sensor manufacturers — some output pressure distribution matrices, some output 3D force vectors, some output vibration frequencies. Everyone's data is incompatible. Research results can't be reproduced. The whole field feels like everyone is fighting their own battle.
Third, touch and action are too tightly coupled. Vision can be done “standing still and watching,” but touch requires “reaching out and touching.” If you want to collect tactile data, you have to make the robot physically interact with objects — which means data collection is N times more expensive than vision, and N times riskier.
The Meaning of 30,000 Hours: More Than Data — It's a Standard
The most valuable aspect of the three reports from Xinzhi Embodied and Fudan isn't the 30,000-hour number — it's that they've established a unified tactile data format and collection specification.
This is a bit like what ImageNet meant for computer vision. What made ImageNet valuable wasn't the 14 million images — it was that it established a unified benchmark. Everyone trains on the same data and evaluates by the same standards, so research can accelerate along the same track.
The release of the 30,000-hour tactile dataset could become the “ImageNet moment” of embodied AI. With a unified data format and benchmark, research results from different teams can be compared, reproduced, and built upon each other. More importantly, with large-scale real tactile data, robots can truly learn “how to apply force” — how hard to squeeze an egg, how much torque to use turning a screw, how to hold a cup level, how to control force variation when kneading dough.
At WAIC 2026, embodied AI and intelligent computing were listed as the two core tracks, with over 200 companies, 200+ embodied AI terminals, and 300+ physical robots on display. But anyone paying attention can see: most robots are still stuck at the “can walk and talk” stage. Very few can actually “get work done.” The core reason they can't work isn't that their brains aren't smart enough (LLMs are already plenty strong) or that their eyes aren't good enough (vision models are improving fast) — it's that their hands aren't dexterous enough and their sense of touch isn't fine enough.
Until the touch barrier is crossed, embodied AI willforever confined to dancing on stage — it can dance on stage at exhibitions, but it can't enter factories and homes. The 30,000-hour dataset is a beginning, but it marks the point where the industry is finally starting to take this “most embarrassing shortcoming” seriously.
See you tomorrow.
视觉让机器人看得见,大模型让机器人想得到,但只有触觉能让机器人做得到。具身智能的最后一公里,在指尖。
—— Dawn Vision编辑部
Vision lets robots 'see,' LLMs let robots 'think,' but only touch lets robots 'do.' The last mile of embodied AI is at the fingertips.
—— The Dawn Vision Editorial Desk
embodied AI · tactile data · 30000 hours · Xinzhi Embodied · Fudan University · open dataset · tactile perception · robotic dexterous manipulation
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 5 个源信号生成,经编辑部人工审核。素材来源:量子位、今日头条、WAIC 2026。
This article was generated by the Dawn Vision cognitive engine processing 5 source signals, with human editorial review. Sources: QbitAI, Jinri Toutiao, WAIC 2026.