算力基建

SK海力士+闪迪发布HBF标准
AI SSD填补HBM与NVMe空白

SK Hynix + SanDisk Launch HBF Standard
AI SSD Fills Gap Between HBM and NVMe

HBM太贵,SSD太慢——SK海力士与闪迪联手推出HBF高带宽闪存,3TB/s带宽、512GB容量、成本仅HBM的1/8,推理存储新范式来了。

HBM is too expensive, SSD is too slow — SK Hynix and SanDisk team up to launch HBF high-bandwidth flash with 3TB/s bandwidth, 512GB capacity, at 1/8 the cost of HBM. A new inference storage paradigm is here.

No.031 2026.08.07 约 5 分钟阅读 ~5 min read

AI推理的存储瓶颈,可能要被一种新硬件打破了。

8月4日的闪存峰会上,SK海力士(SK Hynix)和闪迪(SanDisk,西部数据旗下)联合发布了一个全新的存储标准——HBF(High Bandwidth Flash,高带宽闪存)。简单说,这是一种介于HBM和普通NVMe SSD之间的新存储层级:比HBM容量大、便宜很多,又比普通SSD快得多。

具体规格:最高512GB容量,带宽最高3TB/s。作为对比,HBM3e的单颗容量通常只有24-36GB,但带宽能到8-10TB/s;普通NVMe SSD容量能到几十TB,但带宽通常只有10-15GB/s。HBF正好卡在中间——容量比HBM大一个数量级以上,带宽比普通SSD高两个数量级。

价格呢?HBM之父金正浩教授有个很形象的比喻:HBF的成本大约是HBM的1/8。考虑到HBM现在是AI服务器里最贵的组件之一(每颗价格数千美元),这个成本差距是数量级的。

NVIDIA也没闲着,同期推出了BlueField-4驱动的CMX(Context Memory Xcelerator)上下文内存存储平台——说白了就是为AI推理专门优化的存储加速方案,能把模型上下文存在靠近计算单元的地方,大幅降低推理延迟和成本。

巨头们同时发力AI推理存储,这件事的信号意义很强。

为什么AI推理需要新的存储层级?

要理解HBF的价值,你得先理解大模型推理的痛点。

现在的大模型推理,最大的瓶颈之一是模型加载和上下文切换。一个70B参数的模型,光权重就有140GB左右;如果是MoE模型,总参数更是动辄几千亿。这些模型存在哪里?以前的方案是:全部放进HBM里——快,但贵到离谱;或者放在SSD里——便宜,但加载慢到不能忍。

这就形成了一个尴尬的断层:HBM太快太贵容量太小,SSD太慢太便宜容量太大。中间缺了一层——既要有接近HBM的速度,又要有SSD级别的容量和成本。

HBF瞄准的就是这个空白。3TB/s的带宽虽然还比不上HBM,但对于很多推理场景来说已经够用了——尤其是那些不需要极致低延迟、但需要大上下文和高吞吐量的场景,比如RAG、长文档处理、批量推理。512GB的单颗容量意味着一两颗HBF就能装下一个大模型的全部权重,不用再从SSD慢慢加载。

更关键的是成本。如果HBF真的只要HBM 1/8的价格,那AI推理服务器的BOM成本结构会被彻底改写。以前一台推理服务器可能要配十几颗HBM,光内存成本就几万美元;未来可能只需要几颗HBM做热点缓存,再配上几颗HBF做主存储,总成本可能下降30-50%。

新存储层级会带来什么变化?

HBF这类AI SSD的普及,可能会从几个层面改变AI行业的格局。

第一,推理成本大幅下降。这是最直接的影响。存储成本降下来了,推理的单位token成本就能降,大模型API的价格还有很大的下降空间——当然,这取决于厂商愿不愿意把成本下降的红利让给用户。

第二,大模型部署门槛降低。以前部署一个70B模型可能需要一台几十万甚至上百万的服务器;有了HBF之后,可能一台几万块的服务器就能跑起来。这意味着更多中小企业、甚至个人开发者都能部署自己的大模型,AI的民主化会更进一步。

第三,长上下文能力成为标配。现在的长上下文模型(百万token级别)之所以贵,很大程度上是因为需要巨大的HBM容量来存KV缓存。如果HBF能以更低的成本提供大容量高速存储,百万token上下文可能就不再是高端模型的专属功能,而是变成所有模型的标配。

当然,HBF也不是没有短板。写入寿命就是一个问题——闪存的写入次数是有限的,而HBM的写入寿命要长得多。对于需要频繁写入的场景(比如训练),HBF可能暂时还顶不上。但对于以读为主的推理场景,HBF的短板影响不大。

总的来说,HBF的出现是AI算力基建的一个重要里程碑。它不是要替代HBM,而是要填补HBM和SSD之间的空白,让AI推理的存储架构更加分层、更加高效。当存储不再是瓶颈的时候,AI推理的成本曲线,可能会迎来一个新的拐点。

明天见。

The storage bottleneck in AI inference might be about to be broken by a new type of hardware.

At the Flash Memory Summit on August 4, SK Hynix and SanDisk (a Western Digital brand) jointly released a brand new storage standard — HBF (High Bandwidth Flash). Simply put, it's a new storage tier between HBM and regular NVMe SSDs: more capacity and much cheaper than HBM, yet much faster than regular SSDs.

Specific specs: up to 512GB capacity, bandwidth up to 3TB/s. For comparison, HBM3e typically only has 24–36GB per stack but reaches 8–10TB/s bandwidth; regular NVMe SSDs can reach tens of terabytes in capacity but only 10–15GB/s bandwidth. HBF sits right in the middle — an order of magnitude more capacity than HBM, and two orders of magnitude higher bandwidth than regular SSDs.

And the price? Professor Kim Jung-ho, the "father of HBM," put it vividly: HBF costs roughly 1/8 of HBM. Considering HBM is currently one of the most expensive components in AI servers (thousands of dollars per stack), this cost gap is orders of magnitude.

NVIDIA isn't sitting idle either — it simultaneously launched the CMX (Context Memory Xcelerator) platform driven by BlueField-4 — essentially a storage acceleration solution specifically optimized for AI inference, keeping model context close to compute units to dramatically reduce inference latency and cost.

Giants simultaneously pushing AI inference storage is a strong signal.

Why Does AI Inference Need a New Storage Tier?

To understand HBF's value, you first have to understand the pain point of LLM inference.

One of the biggest bottlenecks in current LLM inference is model loading and context switching. A 70B-parameter model has ~140GB of weights alone; for MoE models, total parameters can be hundreds of billions. Where do these models live? The old options: put everything in HBM — fast, but ridiculously expensive; or store them on SSD — cheap, but loading is painfully slow.

This creates an awkward gap: HBM is too fast, too expensive, too small; SSD is too slow, too cheap, too big. There's a missing layer in between — something with near-HBM speed but SSD-level capacity and cost.

HBF targets precisely this gap. 3TB/s bandwidth, while not quite matching HBM, is already sufficient for many inference scenarios — especially those that don't need ultra-low latency but require large context and high throughput, like RAG, long-document processing, and batch inference. 512GB per device means one or two HBF modules can hold an entire large model's weights, no more slow loading from SSD.

More critically, cost. If HBF really is only 1/8 the price of HBM, the BOM cost structure of AI inference servers will be completely rewritten. Previously, an inference server might need a dozen HBM stacks, with memory costs alone in the tens of thousands of dollars; in the future, it might only need a few HBM stacks for hot cache plus a few HBF modules for main storage, with total cost potentially dropping 30–50%.

What Changes Will a New Storage Tier Bring?

The adoption of AI SSDs like HBF could reshape the AI industry landscape on several levels.

First, inference costs drop dramatically. This is the most direct impact. Lower storage costs mean lower per-token inference costs, and there's still plenty of room for LLM API prices to come down — though whether vendors pass the savings on to users is another question.

Second, LLM deployment barriers lower. Previously, deploying a 70B model might require a server costing hundreds of thousands of yuan; with HBF, it might run on a server costing tens of thousands. This means more SMEs and even individual developers can deploy their own LLMs, and AI democratization takes another step forward.

Third, long-context capability becomes standard. The reason current long-context models (million-token level) are expensive is largely because they require enormous HBM capacity for KV cache. If HBF can provide large-capacity high-speed storage at lower cost, million-token context might no longer be an exclusive feature of premium models — it becomes standard across all models.

Of course, HBF isn't without its weaknesses. Write endurance is one issue — flash memory has limited write cycles, while HBM's write endurance is much higher. For write-heavy scenarios like training, HBF probably can't replace HBM yet. But for read-heavy inference scenarios, HBF's weaknesses don't matter much.

Overall, the emergence of HBF is an important milestone in AI compute infrastructure. It's not meant to replace HBM — it's meant to fill the gap between HBM and SSD, making AI inference storage architecture more layered and more efficient. When storage is no longer the bottleneck, AI inference's cost curve might reach a new inflection point.

See you tomorrow.

HBF的成本大约是HBM的1/8——AI推理服务器的BOM成本结构,可能要被彻底改写。

—— Dawn Vision 判断

HBF costs roughly 1/8 of HBM — the BOM cost structure of AI inference servers might be completely rewritten.

— Dawn Vision analysis
算力基建 · HBF · AI SSD · SK海力士 · 闪迪 · HBM · NVMe · 推理成本 · 存储架构 · NVIDIA CMX · BlueField-4
Compute Infrastructure · HBF · AI SSD · SK Hynix · SanDisk · HBM · NVMe · Inference Cost · Storage Architecture · NVIDIA CMX · BlueField-4
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的公开信息整理,素材来源:InfoQ中文、闪存峰会、NVIDIA开发者博客、搜狐科技。

This article is based on public information processed by Dawn Vision. Sources: InfoQ Chinese, Flash Memory Summit, NVIDIA Developer Blog, Sohu Tech.