算力基建 · 芯片战争

Jalapeño公布首批基准成绩
InferenceX下每瓦效率达GB300的1.5-1.9倍

Jalapeño Posts First Benchmark Results
1.5-1.9x Per-Watt Workload vs GB300 Under InferenceX

自研芯片启动九个月即出成绩:SemiAnalysis公开基准下,Jalapeño每瓦工作量最高1.9倍于GB300、延迟降低最高3.6倍——而封装功耗只有对手一半。

Nine months after kickoff, OpenAI's custom inference chip posts first results: up to 1.9x per-watt workload versus GB300 and latency cut by up to 3.6x — with half the package power.

No.044 2026.08.27 约 5 分钟阅读 ~5 min read

距项目启动仅约 九个月,OpenAI 的自研推理芯片 Jalapeño 就交出了第一份对外成绩单。这次它进的依然是"别人家的考场":SemiAnalysis 维护的公开推理基准 InferenceX——把自家硬件放进第三方监考的评估体系,正在成为大厂交卷的标准姿势。

成绩单长什么样

官方口径如下:测试覆盖三个模型——GPT-OSS 120B、DeepSeek R1 670B、Kimi K2.5 1T,统一采用 MXFP4 精度、nominal 8k/1k 上下文与 STP 设置。在峰值吞吐场景下,Jalapeño 每瓦完成的 AI 工作量为对比系统 GB200/GB300 的 1.5 至 1.9 倍;端到端延迟方面的优势更为醒目,达到 1.7 至 3.6 倍。而整场对比里最刺眼的数字藏在物理层:Jalapeño 封装 TDP 为 700W,GB300 为 1400W——正好一半。

为什么主打"每瓦"和"延迟"

这两个指标的选定颇富指向性。每瓦工作量对应推理经济学的核心命题:当模型毛利被电费一口口吃掉,能效比就是利润率本身。低批量下的低延迟则瞄准 Agent 工作负载的真实形态——大量短上下文、高频次的小请求,而不是训练式长任务的热核满载。换句话说,Jalapeño 不是要去抢 GPU 的所有跑道,它盯住的是下一代流量结构里的主力泳道

边界必须划清楚

需要把话说严:这是特定负载、固定精度、非对等功率条件下的相对优势,不是对英伟达的全面超越。700W 对 1400W 本就条件不一,官方亦未给出超大规模集群的吞吐对照。若因此宣称"英伟达护城河失守",属于过度解读。

但方向性已经足够清晰:九个月拿出首测结果,说明这条自研管线绝不是 PPT 项目。叠加 AWS、Google、Meta 各自的自研芯片布局,"去英伟达化"正从个别巨头的生存焦虑演变为行业标准动作——而下一次采购谈判桌上谁的语气更平静,将由这些数字决定。

Roughly nine months after the project kicked off, OpenAI's custom inference chip Jalapeño has turned in its first public scorecard. And once again it sat for somebody else's exam: InferenceX, the public inference benchmark maintained by SemiAnalysis — submitting your own silicon to a third-party proctor is fast becoming standard posture when big tech ships numbers.

What the Report Card Looks Like

Per the official write-up: three models were tested — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — uniformly configured with MXFP4 precision, nominal 8k/1k contexts, and STP settings. At peak throughput, Jalapeño completes 1.5 to 1.9 times more AI work per watt than the comparison systems, GB200/GB300. The end-to-end latency gap is even starker: 1.7 to 3.6 times lower. Meanwhile the most eye-catching number sits in the physical layer: Jalapeño's package TDP is 700W against GB300's 1400W — exactly half.

Why "Per-Watt" and "Latency"?

The choice of metrics is pointed. Work per watt addresses inference economics head-on: once electricity eats through model gross margins, energy efficiency is the margin. Low latency at small batch sizes targets the real shape of agentic workloads — torrents of short-context, high-frequency requests rather than long training-style saturation runs. In other words, Jalapeño isn't contesting every lane GPUs occupy; it has fixed its sights on the main swimming lane of the next traffic structure.

Drawing the Boundary Clearly

Say it strictly: this is a relative advantage under specific workloads, fixed precision, and non-equivalent power conditions — not a wholesale defeat of NVIDIA. A 700W-versus-1400W comparison is uneven to begin with, and no hyperscale cluster throughput data has been offered. Reading it as "NVIDIA's moat has fallen" would be overreach.

Still, the direction is unambiguous: a first set of results nine months in means this custom pipeline is no slideware project. Layer on the in-house silicon programs at AWS, Google, and Meta, and de-NVIDIA-fication is mutating from one giant's survival anxiety into an industry-standard move — and the calmest voice at the next procurement negotiation will be decided by exactly these digits.

芯片战争的胜负不会写在发布会的掌声里,而会写在每一度电换来的 token 上。

—— Dawn Vision编辑部

The chip war won't be decided in the applause of a launch event, but in every token bought per kilowatt-hour.

— The Dawn Vision Editorial Desk
Jalapeño · InferenceX公开基准 · 每瓦工作量1.5-1.9倍 · 延迟降低1.7-3.6倍 · TDP 700W vs 1400W · 九个月出首测 · 特定负载口径
Jalapeño · InferenceX public benchmark · 1.5-1.9x per-watt workload · 1.7-3.6x lower latency · TDP 700W vs 1400W · First results in ~9 months · Specific-workload caveat
Sources · 信源 Sources

本文基于 OpenAI 官方博客一手信息,辅以 The Verge 报道交叉确认;所有对比数据限定于 SemiAnalysis InferenceX 公开基准下的指定工作负载与精度配置,功率条件不对等已在正文标注。

Built on OpenAI's official blog as the primary source, cross-confirmed with The Verge coverage; all comparative figures are scoped to designated workloads and precision configs on the SemiAnalysis InferenceX public benchmark, with the unequal power conditions flagged in the text.