64张卡。一台机器。
8月12日,阿里云灵骏真武M890超节点实例在乌兰察布正式开售。这个产品最硬核的一个数字是:单实例成功运行超2万亿参数的大模型——这是国内首个做到这一点的超节点形态算力,Kimi K3和Qwen3.8 Max都已经跑在上面对外服务。
过去,跑一个万亿参数的大模型,企业要么自建机房买几千张卡,要么在海外高端GPU上排队。现在,阿里云把64张自研芯片打包成一个“开箱即用”的AI超级计算机,按需开通即可。
这件事的意义,远不止一个云产品的上新。它标志着国产算力的竞争逻辑,正在发生一次根本性的转向:从“拼单卡性能”,转向“拼系统能力”。
超节点:用系统弥补单卡的代差
先看几个关键参数。
真武M890超节点通过自研的ICN Switch 1.0交换芯片,把Scale-up互联规模从16卡拉升到64卡,卡间互联带宽提升到800GB/s;单卡144GB显存,64卡组成9216GB的显存池。
客观地说,单颗真武M890芯片的算力,距离海外顶级GPU仍有代差。但阿里云的选择是:不在单卡上硬碰硬,用高密度互联和系统级优化,把64张卡拧成一股绳。
结果就是:在智能驾驶、具身智能等训练场景,M890超节点相比上一代810E,训练性能提升了3倍;推理端最高可承载十万亿参数级的MoE大模型。
这套思路,和华为昇腾超节点、百度昆仑芯超节点走的是同一条路——既然单颗芯片追不上,那就让一堆芯片协同,用系统架构抹平单卡差距。
从“实验室样机”到“开箱即用”,门槛被拉低了多少
更关键的变化,在商业落地。
过去国产超节点大多是政企定制机房,投入动辄数亿,中小企业根本玩不起。这次阿里云直接把它做成云上实例,客户不用自建机房,按需开通64卡算力单元就行。
门槛的下降是量级性的。以前想训练或推理万亿参数大模型,得自建机房、拉光纤、做调度,没有几个亿资金根本转不动;现在在云上开个实例就能跑起来,AI创业的算力门槛,从“亿级”降到了“百万级”。
“国产算力真正的破局,不是造出比英伟达更强的单卡,而是用系统架构把一堆国产芯片拧成一台能跑万亿参数的超级计算机。这条路,现在走通了。”—— 一位AI基础设施从业者
乌兰察布这个首发地域也很有意思:这座曾经以土豆闻名的城市,如今被业内称为“token工厂”——绿电占比约90%,气候凉爽,正适合高密度算力集群。阿里云把超节点放在这里,本身就是在为大规模国产算力扩容铺路。
当国产超节点从实验室样机进入大规模商用,整条产业链都会被重新定价:上游的国产芯片、交换芯片、液冷服务器,中游的智算中心、高速光模块,下游的大模型厂商、自动驾驶公司——每一个环节,都在国产算力的这条新路径上被重新激活。
单卡的时代正在过去,系统的时代刚刚开始。
明天见。
64 cards. One machine.
On August 12, Alibaba Cloud officially opened sales of its Lingjun Zhenwu M890 supernode instance in Ulanqab. The product's hardest number: a single instance successfully runs large models exceeding 2 trillion parameters - the first supernode-form compute in China to pull this off, with both Kimi K3 and Qwen3.8 Max already served on it.
Before, running a trillion-parameter model meant either building your own data center with thousands of cards or waiting in line for overseas high-end GPUs. Now Alibaba Cloud bundles 64 self-developed chips into a "plug-and-play" AI supercomputer, provisioned on demand.
This is far more than a new cloud product. It marks a fundamental pivot in China's compute logic: from competing on single-chip performance to competing on system capability.
The Supernode: Closing the Chip Gap With Systems
Look at a few key specs first.
The Zhenwu M890 supernode uses a self-developed ICN Switch 1.0 chip to scale up interconnect from 16 cards to 64, lifting inter-card bandwidth to 800GB/s; 144GB of HBM per card, 9,216GB across 64 cards.
Objectively, a single Zhenwu M890 chip still lags overseas top-tier GPUs. But Alibaba Cloud's choice is: don't fight chip-to-chip - use high-density interconnect and system-level optimization to weave 64 cards into one rope.
The result: in autonomous driving and embodied intelligence training, the M890 supernode delivers 3x the training performance of the previous-gen 810E; on inference it can handle MoE models up to 10 trillion parameters.
This is the same path as Huawei's Ascend supernodes and Baidu's Kunlun-chip supernodes - if a single chip can't catch up, make a cluster of them cooperate and erase the gap through system architecture.
From Lab Prototype to Plug-and-Play: How Much Did the Bar Drop
The bigger change is commercial adoption.
Domestic supernodes used to be custom-built for government and enterprise data centers, costing hundreds of millions - out of reach for SMBs. This time Alibaba Cloud turned it into a cloud instance: customers don't build their own data centers, they just provision 64-card compute units on demand.
The drop is order-of-magnitude. Training or inferring a trillion-parameter model used to require building a data center, laying fiber, and engineering scheduling - nothing moves without several hundred million yuan. Now you spin up an instance in the cloud, and the compute barrier for AI startups falls from "hundreds of millions" to "millions."
"China's real compute breakthrough isn't building a single card stronger than Nvidia's - it's using system architecture to weave a stack of domestic chips into a supercomputer that runs trillion-parameter models. That path is now open." - An AI Infrastructure Practitioner
The Ulanqab debut location is telling too: a city once known for potatoes is now called the "token factory" by the industry - about 90% green energy, cool climate, ideal for high-density compute. Placing the supernode here is itself groundwork for scaling domestic compute.
As domestic supernodes move from lab prototypes into large-scale commercial use, the whole supply chain gets repriced: upstream domestic chips, switch chips, liquid-cooled servers; midstream AI data centers and high-speed optical modules; downstream LLM vendors and autonomous-driving companies - every link gets re-energized along this new domestic-compute path.
The single-card era is fading. The systems era is just beginning.
See you tomorrow.
国产算力真正的破局,不是造出比英伟达更强的单卡,而是用系统架构把一堆国产芯片拧成一台能跑万亿参数的超级计算机。这条路,现在走通了。
—— 一位AI基础设施从业者
China's real compute breakthrough isn't building a single card stronger than Nvidia's - it's using system architecture to weave a stack of domestic chips into a supercomputer that runs trillion-parameter models. That path is now open.
- An AI Infrastructure Practitioner
Alibaba Cloud,Zhenwu M890,supernode,China compute,2T parameters,Ulanqab,system-level innovation
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 9 个源信号生成,经编辑部人工审核。素材来源:阿里云官方、环球网、Global Times。
Generated by the Dawn Vision cognitive engine processing 9 source signals, with human editorial review. Sources: Alibaba Cloud Official, Huanqiu.com, Global Times.