Focus · 焦点

OpenAI 的
Jalapeño

OpenAI's
Jalapeño

9个月流片、推理成本砍半、功耗降三成——但博通要微软先买40%才肯开工,资本市场用脚投票。"模型定义芯片"的时代真的来了吗?

9 months from design to tape-out, inference costs cut in half, power consumption down 30% -- but Broadcom wants Microsoft to commit to buying 40% before starting production, and capital markets are voting with their feet. Has the era of 'models defining chips' truly arrived?

No.002 2026.06.25 约 9 分钟阅读 ~9 min read

北京时间 6 月 25 日,OpenAI 联合博通正式发布了首款定制 AI 推理芯片——Jalapeño(墨西哥辣椒)。

从立项到流片仅 9 个月,刷新了高性能 ASIC 的开发周期纪录。要知道,这个行业通常需要 18-24 个月。博通 CEO 陈福阳在发布会上宣称,这颗芯片比传统 GPU 节省约 50% 推理成本、功耗降低约 30%,工程样品已经在实验室里稳定跑着 GPT-5.3-Codex-Spark。

听起来很美好。但资本市场的反应耐人寻味:消息发布当天,博通股价仅涨 1.6%。

要知道,2025 年双方签约的消息传出时,博通当天涨了 15%。这 1.6% 的涨幅,更像是一种礼貌性的鼓掌——"挺好的,然后呢?"

辣椒很辣,但谁来买单?

核心问题悬而未决:钱。

首阶段 1.3 吉瓦产能的芯片生产成本约 180 亿美元。博通的态度很明确:微软得先承诺采购 40%,否则融资不启动。微软呢?至今没有点头。

这不是一笔小账。OpenAI 自己预测,到 2029 年运营烧钱将超过 2000 亿美元。与此同时,他们还在推进 5000 亿美元的 Stargate 超级数据中心项目。一颗 Jalapeño 只是开胃菜,后面的满汉全席才是真正的账单。

"芯片行业有句老话:设计一颗芯片不难,难的是让别人为它买单。" —— 半导体行业观察者

理解这背后的博弈,需要回到一个更本质的问题:为什么 AI 公司突然开始自己做芯片了?

从"用芯片"到"定义芯片":范式正在转移

过去三十年,芯片行业的逻辑是"芯片定义软件"——英特尔造出更快的 CPU,微软和开发者们再想办法把这些算力用起来。英伟达造出更强的 GPU,AI 实验室再围绕 CUDA 生态构建模型。

但 2026 年的今天,这个逻辑正在反转。

大模型的运算模式高度结构化、高度可预测:大量的矩阵乘法、固定的注意力机制、规律的内存访问模式。当你知道自己要跑什么 workload,通用 GPU 的很多设计就变成了浪费——就像你明知道每天只走同一条路上班,却非要开一辆全地形越野车。

于是"模型定义芯片"成为可能:AI 公司比任何人都清楚自己的模型需要什么样的硬件,他们可以反过来设计芯片——去掉不需要的电路,优化关键路径,为 Transformer 架构量身定制计算单元。

Jalapeño 不是第一个吃螃蟹的。谷歌 TPU 已经走了 13 年,亚马逊 Trainium/Inferentia 也迭代到了第三代,微软的 Athena、Meta 的 MTIA 都在推进中。但 OpenAI 的特殊之处在于:它是第一个真正意义上"模型公司"做芯片——它不卖云服务,不做广告,它的全部收入和未来都押注在模型本身。

这形成了一个诱人的飞轮:AI 设计芯片 → 芯片运行 AI → AI 变强 → 设计下一代芯片。理论上,这个飞轮一旦转起来,英伟达的 GPU 霸权将被釜底抽薪。

但飞轮转起来之前,你得先推第一把

理论很美,现实很骨感。

谷歌 TPU 走了 13 年,至今谷歌云仍然在大量采购英伟达 GPU。为什么?因为自研芯片的容错空间太小了。

英伟达的 CUDA 生态花了二十年构建,数百万开发者、数不清的软件库、成熟的工具链——这不是"芯片快 50%"就能轻易替代的。而且 GPU 是通用的:今天跑大模型,明天可以跑物理仿真,后天可以渲染电影。但 ASIC 是为特定 workload 定制的,如果模型架构变了呢?

想想看:2023 年大家都在堆稠密 Transformer,2024 年 MoE 成为主流,2025 年状态空间模型异军突起,2026 年的今天谁也说不清下一代架构是什么。一颗从设计到量产需要 18 个月的芯片,面对的是一个每 6 个月就可能颠覆自己的行业。

这就是 Jalapeño 面临的悖论:它为 GPT 架构优化,但如果 GPT-6 采用了完全不同的架构怎么办?这 180 亿美元的投入会不会变成下一个 Itanium?

更微妙的是 OpenAI 和微软的关系。微软是 OpenAI 最大的投资者、最大的云服务商、最重要的分销渠道。但现在 OpenAI 要自己做芯片,本质上是在减少对英伟达的依赖,也是在减少对 Azure 的依赖——微软不可能看不透这一点。

"你要我花 70 多亿美元买你的芯片,来削弱我自己的云业务竞争力?"——微软的犹豫完全可以理解。

真正的信号:不是芯片,是控制权

但 Jalapeño 的意义,可能远不止于一颗芯片本身。

2026 年的 AI 行业正在进入一个新阶段:模型能力的差距在缩小,但基础设施的控制力在分化。谁掌握了算力,谁就掌握了模型迭代的速度;谁掌握了芯片,谁就掌握了算力的成本曲线。

OpenAI 不是在和英伟达竞争,它是在争夺自己命运的控制权。如果推理成本不能持续下降,Agent 时代的大规模商用就永远是 PPT 上的故事——你不可能让每个用户都花 20 美元/月来供养一个需要昂贵 GPU 才能运行的 Agent。

50% 的成本下降意味着什么?意味着原本需要 100 亿美元算力才能支撑的产品,现在只需要 50 亿。意味着 Agent 从"高端玩具"变成"大众工具"的临界点,可能因为这一颗芯片而提前到来。

这也是为什么资本市场虽然反应冷淡,但没有人敢真正看空。大家都在等一个信号:微软会不会签那张 70 亿美元的采购单。

签了,飞轮启动,AI 行业的成本曲线将被重新定义。不签,Jalapeño 可能就是另一颗耀眼但短暂的流星——就像过去十年里无数诞生于发布会、消失于量产线的"芯片颠覆者"一样。

辣椒入菜,还需时日

Jalapeño 发布后的 24 小时,AI 圈的讨论已经从"芯片性能"转向了"商业可行性"。这本身就是一个信号:AI 硬件正在从"技术叙事"进入"商业叙事"阶段。

9 个月流片确实是工程奇迹,成本减半确实足够诱人,功耗降低确实是数据中心运营商梦寐以求的。但半导体行业从来不缺奇迹,缺的是能持续赚钱的奇迹。

对于普通从业者和观察者来说,Jalapeño 真正值得关注的不是参数,而是时间线:如果一切顺利,这颗芯片将在 2027 年量产,2028 年大规模部署。那时候,GPT-6 可能已经发布,Agent 可能已经成为主流交互方式,AI 的竞争格局可能已经完全不同。

但无论如何,一个趋势已经清晰:AI 公司正在向上游延伸,从模型到芯片到数据中心,垂直整合正在成为头部玩家的标配。就像特斯拉不满足于买车、还要造电池、还要做自动驾驶芯片一样,OpenAI 也在构建自己的垂直帝国。

墨西哥辣椒很辣,但能不能辣到英伟达嘴里,还要看微软愿不愿意当这个"第一个吃辣椒的人"。


明天见。

On June 25 Beijing time, OpenAI and Broadcom officially unveiled the first custom AI inference chip -- Jalapeño.

From project initiation to tape-out took just 9 months, breaking the development cycle record for high-performance ASICs. For context, this industry typically requires 18-24 months. Broadcom CEO Hock Tan claimed at the launch that this chip saves roughly 50% on inference costs versus traditional GPUs and reduces power consumption by about 30%, with engineering samples already running GPT-5.3-Codex-Spark stably in the lab.

It sounds great. But capital markets' reaction was telling: on the day of the announcement, Broadcom stock rose just 1.6%.

For context, when news of the partnership broke in 2025, Broadcom rose 15% that day. That 1.6% gain felt more like a polite round of applause -- "nice, now what?"

The Pepper Is Hot, But Who Pays?

The core question remains unresolved: money.

First-phase production costs for 1.3 GW of chip capacity are roughly $18 billion. Broadcom's position is clear: Microsoft must commit to purchasing 40% first, or financing doesn't start. And Microsoft? Hasn't nodded yet.

This isn't a small bill. OpenAI itself projects operational burn will exceed $200 billion by 2029. Meanwhile, they're pushing the $500 billion Stargate super-datacenter project. One Jalapeño is just the appetizer; the real bill is the full-course meal behind it.

"There's an old saying in the chip industry: designing a chip isn't hard; getting someone else to pay for it is." -- Semiconductor industry observer

To understand the game behind this, you need to return to a more fundamental question: why are AI companies suddenly making their own chips?

From "Using Chips" to "Defining Chips": The Paradigm Is Shifting

For the past thirty years, the chip industry logic was "chips define software" -- Intel built faster CPUs, and Microsoft and developers figured out how to use that compute. Nvidia built stronger GPUs, and AI labs built models around the CUDA ecosystem.

But today in 2026, that logic is reversing.

LLM computation patterns are highly structured and highly predictable: massive matrix multiplications, fixed attention mechanisms, regular memory access patterns. When you know exactly what workloads you're running, much of a general-purpose GPU's design becomes waste -- like knowing you only drive the same route to work every day but insisting on an all-terrain off-road vehicle.

Thus "models defining chips" becomes possible: AI companies know better than anyone what hardware their models need; they can reverse-engineer chips -- strip unnecessary circuits, optimize critical paths, build compute units tailor-made for the Transformer architecture.

Jalapeño isn't the first to the party. Google TPU has been at it for 13 years; Amazon Trainium/Inferentia is on its third generation; Microsoft's Athena and Meta's MTIA are both in progress. But what makes OpenAI special is: it's the first true "model company" making chips -- it doesn't sell cloud services, doesn't do advertising; its entire revenue and future are bet on the model itself.

This creates an enticing flywheel: AI designs chips → chips run AI → AI gets stronger → designs next-generation chips. Theoretically, once this flywheel spins, Nvidia's GPU hegemony would be pulled out from under it.

But Before the Flywheel Spins, You Have to Give It the First Push

Theory is beautiful; reality is brutal.

Google TPU has been going for 13 years, and Google Cloud still buys massive amounts of Nvidia GPUs. Why? Because custom chips have too little room for error.

Nvidia's CUDA ecosystem took twenty years to build -- millions of developers, countless software libraries, mature toolchains -- this isn't something "50% faster chips" can easily replace. And GPUs are general-purpose: run LLMs today, physics simulations tomorrow, movie rendering the day after. But ASICs are customized for specific workloads; what if the model architecture changes?

Think about it: in 2023 everyone was stacking dense Transformers; in 2024 MoE went mainstream; in 2025 state-space models emerged out of nowhere; in 2026 today nobody can say what the next architecture is. A chip that takes 18 months from design to mass production faces an industry that might upend itself every 6 months.

That's the paradox Jalapeño faces: it's optimized for GPT architecture, but what if GPT-6 uses a completely different architecture? Could this $18 billion investment become the next Itanium?

More subtly, there's the OpenAI-Microsoft relationship. Microsoft is OpenAI's largest investor, largest cloud provider, most important distribution channel. But now OpenAI making its own chips is essentially reducing dependence on Nvidia -- and also reducing dependence on Azure -- Microsoft can't fail to see this.

"You want me to spend $7+ billion buying your chips to weaken my own cloud business competitiveness?" -- Microsoft's hesitation is entirely understandable.

The Real Signal: It's Not About Chips, It's About Control

But Jalapeño's significance may go far beyond a single chip.

The AI industry in 2026 is entering a new phase: gaps in model capability are narrowing, but control over infrastructure is diverging. Who controls compute controls the speed of model iteration; who controls chips controls the compute cost curve.

OpenAI isn't competing with Nvidia; it's fighting for control over its own destiny. If inference costs can't keep falling, large-scale Agent commercialization will forever be a story on PPT slides -- you can't have every user spending $20/month to sustain an Agent that needs expensive GPUs to run.

What does a 50% cost drop mean? It means a product that originally needed $10 billion in compute to support now only needs $5 billion. It means the tipping point where Agents go from "premium toy" to "mass-market tool" might arrive early because of this one chip.

That's also why capital markets, though reacting coolly, haven't dared to truly go short. Everyone is waiting for one signal: will Microsoft sign that $7 billion purchase order?

If they sign, the flywheel starts, and the AI industry's cost curve gets redefined. If they don't, Jalapeño might be another bright but brief meteor -- like the countless "chip disruptors" over the past decade that were born at launch events and died on production lines.

Putting the Pepper in the Dish Takes Time

In the 24 hours after Jalapeño's launch, AI circle discussion shifted from "chip performance" to "commercial viability." That in itself is a signal: AI hardware is moving from the "technology narrative" phase into the "commercial narrative" phase.

9-month tape-out is indeed an engineering miracle; halving costs is tantalizing enough; reduced power consumption is the stuff of datacenter operators' dreams. But the semiconductor industry has never lacked miracles; what it lacks are miracles that sustainably make money.

For ordinary practitioners and observers, what Jalapeño really deserves attention for isn't specs, but the timeline: if everything goes smoothly, this chip will enter mass production in 2027, with large-scale deployment in 2028. By then, GPT-6 might have launched, Agents might be the mainstream interaction mode, and the AI competitive landscape might be entirely different.

But regardless, one trend is clear: AI companies are extending upstream, from models to chips to datacenters; vertical integration is becoming standard for top players. Just as Tesla wasn't content selling cars and also had to build batteries and make autonomous driving chips, OpenAI too is building its own vertical empire.

The jalapeño is spicy, but whether it burns Nvidia's mouth depends on whether Microsoft is willing to be the "first to eat the pepper."


See you tomorrow.

芯片行业有句老话:设计一颗芯片不难,难的是让别人为它买单。

—— 半导体行业观察者
AI 芯片自研潮全景 · 谷歌 TPU 十三年启示录 · 英伟达 CUDA 护城河解析 · OpenAI 垂直整合战略
AI custom chip panorama, Google TPU 13-year lessons, Nvidia CUDA moat analysis, OpenAI vertical integration strategy
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 11 个源信号自动生成,经编辑部人工审核。素材来源包括:OpenAI Jalapeño 发布会信息、博通财报分析、半导体行业研报、微软采购动态、Stargate 项目追踪。

Auto-generated by Dawn Vision's cognitive engine from 11 source signals, editorially reviewed. Sources include: OpenAI Jalapeño launch info, Broadcom financial analysis, semiconductor industry research, Microsoft procurement dynamics, Stargate project tracking.