110亿美元。
7月8日,AI芯片公司SambaNova宣布完成10亿美元F轮融资的首次交割,估值达到110亿美元。General Atlantic领投,BlackRock、Capital Group、卡塔尔投资局、Vista Equity Partners等全明星阵容跟投,Intel继续加码——5个月前的2月份,SambaNova刚融完3.5亿美元E轮并发布SN50芯片。
比融资更值得注意的是客户:摩根大通(JPMorgan Chase)选定SambaNova作为其推理基础设施合作伙伴,SN40L和SN50系统将为这家全球最大银行提供安全的本地化AI推理服务。SambaNova CEO Rodrigo Liang直言不讳:"摩根大通这样级别的银行决定用SambaNova做推理,这向整个银行业传递了一个信号——是时候不完全依赖云服务了。"
同一天,两条来自中国的消息形成了完美呼应:InfoQ报道DeepSeek一年前已启动自研AI推理芯片项目,正在对接代工和存储厂商;36氪消息称海光信息将集中展示"云边端"全系算力场景,首次大规模切入端侧AI。
训练靠英伟达、推理上公有云——这个AI行业默认了三年的算力格局,正在2026年夏天被彻底改写。
为什么摩根大通选择"弃云从私"?
SambaNova拿到摩根大通这个客户的分量,怎么强调都不为过。
银行业是对AI需求最迫切、同时对安全要求最严苛的行业。摩根大通管理着超过3.9万亿美元的资产,每天处理的交易笔数以亿计,客户数据的敏感度是最高级别。过去两年,摩根大通和其他华尔街大行一样,主要通过AWS、Azure、Google Cloud三大云厂商调用AI能力——OpenAI服务在Azure上,Anthropic在AWS和GCP上都有部署。
但大银行正在集体转向私有部署。Liang提到SambaNova的三类客户——主权云、新云厂商(neocloud)、自建自用的大型企业——摩根大通代表的是第三类,也是最有风向标意义的一类。原因有三:
第一,数据安全和合规。银行最敏感的风控模型、反洗钱模型、客户数据,即使在VPC隔离的云上,很多CIO和CISO依然睡不着觉。本地化部署意味着数据永远不出银行自己的数据中心,满足最严格的监管要求。欧洲的GDPR、美国的银行保密法、中国的数据安全法,都在推动金融机构把核心AI能力迁回本地。
第二,推理成本的非线性增长。训练是一次性投入,推理是持续的成本中心。2025年AI行业的注意力还在训练大模型需要多少钱,2026年大家开始算一笔账:当大模型真正部署到生产环境,每天几亿次调用,推理成本会随着调用量线性甚至超线性增长。云厂商的推理定价虽然在下降,但对超大规模用户来说,自建私有基础设施的TCO(总拥有成本)在2-3年维度上已经显著低于公有云。
第三,异构计算的性能优势。SambaNova的SN50芯片是专门为Agentic Inference(智能体推理)设计的——不是为了训练万亿参数模型,而是为了让万亿参数模型在推理时跑得更快、延迟更低、单token成本更便宜。Liang说SambaNova的核心竞争力是"premium inference":把多万亿参数模型装到一个机架里,让它们跑得快。这和GPU的通用架构思路完全不同——GPU是瑞士军刀,什么都能干;SambaNova的RDU(可重构数据流单元)是手术刀,专门优化推理。
"2024年银行问'我们该不该用AI';2025年银行问'我们该用哪家云厂商的AI';2026年银行问'我们该买谁的芯片建自己的AI集群'。问题的变化本身就是答案。"—— 一位华尔街金融科技分析师
中国侧的推理芯片竞赛:DeepSeek下海、海光端侧突围
SambaNova在北美拿下摩根大通的同一天,中国AI推理芯片赛道也传来两个重要信号。
第一个是DeepSeek被曝自研推理芯片。InfoQ的报道称,DeepSeek大约在一年前就启动了AI推理芯片项目,目前正在对接代工厂和存储厂商。这个消息一点都不让人意外。DeepSeek从V3到R1再到最新的模型,一直走的是极致性价比路线——开源、便宜、推理成本低。但"便宜"是建立在英伟达GPU之上的,当你的调用量达到一定规模,GPU本身的采购成本和功耗就成了最大的成本项。如果能自研推理芯片,DeepSeek就能把推理成本的话语权完全掌握在自己手里。
这和Anthropic找三星造芯、OpenAI做Jalapeño芯片是同一个逻辑。当AI公司的规模做到一定程度,不做芯片就等于把利润的咽喉交给英伟达和台积电。但DeepSeek的路径可能更激进:Anthropic和OpenAI找的是三星这样的成熟代工合作,DeepSeek作为一家中国公司,面临的代工选择和供应链约束要复杂得多。
第二个是海光信息切入端侧AI。36氪报道称,海光将在光合组织2026智能计算应用大会上集中展示"云边端"全系算力场景,这是海光在CPU和DCU(深度学习处理器)双芯底座基础上,首次大规模呈现面向工业现场和行业应用的端侧算力布局。端侧AI的核心需求是实时响应、本地处理、安全可信——这恰恰是很多工业场景、边缘计算场景不愿意用云的原因。
韩国AI芯片公司Rebellions也传出消息,计划2027年第一或第二季度在韩国IPO,目前已经开始产生实际营收。Rebellions背后有三星电子的支持,走的也是专用推理芯片路线。从美国的SambaNova、Etched,到中国的DeepSeek自研、海光全栈,再到韩国的Rebellions,全球AI推理芯片赛道正在从英伟达一家独大,走向多地区、多架构、多场景的百花齐放。
推理时代的算力格局:从"一超多强"到"异构战国"
我们正在经历AI算力格局的一次根本性重构。
训练时代(2020-2025)的逻辑很简单:Scaling Law决定了谁能买到最多的H100/H200/B100,谁就能训练出最强的模型。英伟达是这个时代的绝对王者——CUDA生态壁垒、台积电先进制程优先权、NVLink高速互联,让竞争对手几乎没有机会。AMD的MI300系列虽然性能不错,但生态差距太大;Google TPU只给自己用;各种AI芯片创业公司大多停留在PPT阶段。
但推理时代的游戏规则完全不同。第一,推理不需要最先进的制程。训练需要极限算力密度和显存带宽,推理对制程的要求低得多,成熟制程也能跑出有竞争力的性能/功耗比。这给了更多玩家入场的机会。第二,推理场景高度碎片化。云端推理、私有数据中心推理、边缘推理、端侧推理——每个场景的延迟、功耗、成本要求都不一样,没有一款芯片能通吃所有场景。第三,软件栈的壁垒在降低。PyTorch、ONNX Runtime、vLLM、TensorRT-LLM这些推理框架正在快速成熟,让新芯片的软件适配成本大幅下降。
SambaNova的SN50选择和Intel Xeon深度合作走异构路线,而不是纯自研独立芯片,就是这个趋势的体现。Rodrigo Liang在采访中反复强调和Intel的合作关系——借助Intel的量产能力、供应链规模和客户渠道,SambaNova不需要自己从零开始造一整套系统。Intel也需要SambaNova——Xeon CPU在AI推理上需要专用加速芯片的加持,才能在和NVIDIA Grace Hopper的竞争中不落下风。
微软最近也被曝出加入AI成本削减趋势,更多依赖自研模型。TechCrunch的报道称微软正在减少对OpenAI模型的调用比例,转而使用自研的Phi系列和MaaS模型。这传递的信号同样明确:连OpenAI最大的投资者都在为推理成本做多元化布局,没有人愿意被单一供应商锁死。
终局判断:推理基建的"三层蛋糕"正在形成
把这些线索串在一起,AI推理基础设施的终局轮廓已经清晰可见:一个三层蛋糕结构。
第一层是超大规模公有云推理。AWS、Azure、GCP、阿里云、腾讯云这些云厂商会继续提供通用推理服务,服务对象是中小客户、弹性需求、非敏感场景。这一层的核心竞争力是规模效应和易用性,英伟达GPU在这一层依然会占据主导地位,但份额会从接近垄断降到50-60%。
第二层是行业/主权私有推理基础设施。银行、政府、医疗、大型制造企业会越来越多地建设自己的私有AI推理集群,SambaNova、Etched、Rebellions这类专用推理芯片公司,以及海光、华为昇腾这类国产算力厂商,会在这一层和英伟达正面竞争。这一层的核心竞争力是安全可控、TCO优势和特定场景性能优化。摩根大通选择SambaNova,就是这一层市场爆发的发令枪。
第三层是端侧和边缘推理。手机、PC、汽车、工业设备、机器人——这些场景需要极低延迟、极高隐私性、断网可用,模型会越来越多地跑在端侧而非云端。高通、联发科、苹果、海光、地平线这些玩家会在这一层激烈竞争。
对AI行业的从业者和投资者来说,这个格局变化释放了几个明确信号。第一,推理芯片是比训练芯片更大的市场——训练市场一年几百亿美元到头,推理市场是几千亿甚至上万亿美元的规模,而且持续时间更长。第二,"全自研"不是唯一路径——SambaNova和Intel的合作证明,异构计算、强强联合可能比垂直整合更高效。第三,中国推理芯片厂商的窗口期正在打开——DeepSeek自研、海光全栈布局,不是为了超越英伟达做训练,而是在推理这个更大的市场里抢占国产替代的份额。
110亿美元估值只是开始。AI推理基建的独立战争,才刚刚打响第一枪。
明天见。
$11 billion.
On July 8, AI chip company SambaNova announced the first close of its $1 billion Series F round at an $11 billion valuation. General Atlantic led the round, with an all-star cast including BlackRock, Capital Group, Qatar Investment Authority, and Vista Equity Partners participating, and Intel doubling down — this coming just five months after SambaNova closed a $350 million Series E and launched its SN50 chip in February.
More notable than the funding was the customer: JPMorgan Chase selected SambaNova as an inference infrastructure partner, with SN40L and SN50 systems set to power secure on-premises AI inference at the world's largest bank. SambaNova CEO Rodrigo Liang didn't mince words: "Having JPMorgan Chase decide they're going to use SambaNova for their inference solution is a big deal. It sends a message to the banking industry that it's time not to completely depend on cloud services."
That same day, two pieces of news from China formed a perfect echo. InfoQ reported that DeepSeek had started developing its own AI inference chips about a year ago and was currently engaging foundries and storage vendors. A 36Kr dispatch said Hygon Information would showcase its full-stack "cloud-edge-device" compute portfolio, making a major push into edge AI for the first time.
Train on NVIDIA, inference on the public cloud — this default AI infrastructure playbook that held for three years is being completely rewritten in the summer of 2026.
Why Is JPMorgan Ditching Cloud for On-Prem?
The significance of SambaNova winning JPMorgan cannot be overstated.
Banking is the industry with the most urgent AI needs and the strictest security requirements. JPMorgan manages over $3.9 trillion in assets, processes hundreds of millions of transactions daily, and customer data sensitivity is at the highest possible level. For the past two years, JPMorgan, like other Wall Street giants, primarily accessed AI capabilities through the big three cloud providers — AWS, Azure, and Google Cloud — with OpenAI on Azure and Anthropic deployed across AWS and GCP.
But big banks are collectively shifting to private deployments. Liang identified three categories of SambaNova customers — sovereign clouds, neoclouds, and large enterprises building for their own use — with JPMorgan representing the third and most indicative category. The reasons are threefold:
First, data security and compliance. For banks' most sensitive risk models, anti-money-laundering systems, and customer data, many CIOs and CISOs still can't sleep easy even with VPC-isolated cloud. On-premises deployment means data never leaves the bank's own data centers, satisfying the strictest regulatory requirements. Europe's GDPR, U.S. bank secrecy laws, and China's Data Security Law are all pushing financial institutions to bring core AI capabilities back on-site.
Second, the non-linear growth of inference costs. Training is a one-time investment; inference is a continuous cost center. In 2025, the industry was still focused on how much it cost to train large models; in 2026, everyone's doing the math: when models are truly deployed in production with hundreds of millions of calls per day, inference costs grow linearly — even superlinearly — with usage. While cloud inference pricing keeps falling, for hyperscale users the TCO (total cost of ownership) of building private infrastructure is already significantly lower than public cloud on a 2–3 year horizon.
Third, performance advantages of heterogeneous computing. SambaNova's SN50 chip is purpose-built for Agentic Inference — not for training trillion-parameter models, but for making trillion-parameter models run faster with lower latency and cheaper per-token cost during inference. Liang describes SambaNova's core competitive edge as "premium inference": fitting multi-trillion-parameter models onto a single rack and making them run fast. This is fundamentally different from GPU's general-purpose architecture — GPUs are Swiss Army knives that do everything; SambaNova's RDU (Reconfigurable Dataflow Unit) is a scalpel, optimized specifically for inference.
"In 2024 banks asked 'should we use AI'; in 2025 they asked 'which cloud provider's AI should we use'; in 2026 they're asking 'whose chips should we buy to build our own AI clusters.' The evolution of the question is itself the answer."— A Wall Street fintech analyst
China's Inference Chip Race: DeepSeek Goes Custom, Hygon Pushes Edge
The same day SambaNova won JPMorgan in North America, two important signals emerged from China's AI inference chip race.
First, DeepSeek was reported to be developing its own inference chips. InfoQ's coverage stated that DeepSeek started the AI inference chip project about a year ago and is currently engaging foundry and storage partners. This news shouldn't surprise anyone. From V3 to R1 to its latest models, DeepSeek has consistently pursued extreme cost-efficiency — open source, cheap, low inference cost. But that "cheapness" sits atop NVIDIA GPUs. When your call volume reaches a certain scale, GPU procurement costs and power consumption become the single largest cost line item. By building custom inference chips, DeepSeek can take full control of inference cost economics.
This follows the same logic as Anthropic partnering with Samsung for custom chips and OpenAI building Jalapeño. When AI companies reach a certain scale, not building chips means ceding the throat of your profit margins to NVIDIA and TSMC. But DeepSeek's path may be more aggressive: Anthropic and OpenAI are working with established foundries like Samsung; DeepSeek, as a Chinese company, faces far more complex foundry options and supply chain constraints.
Second, Hygon Information is pushing into edge AI. 36Kr reported that Hygon will showcase its full-stack "cloud-edge-device" compute portfolio at the 2026 Heguang Organization Intelligent Computing Applications Conference, marking Hygon's first large-scale presentation of edge-side compute deployments for industrial and vertical applications, building on its existing CPU and DCU (Deep Learning Processing Unit) dual-chip foundation. Edge AI's core requirements — real-time response, local processing, security and trustworthiness — are precisely why many industrial and edge computing scenarios are reluctant to use the cloud.
Korean AI chip startup Rebellions — backed by Samsung Electronics — also revealed plans to IPO in Korea in Q1 or Q2 2027, and is already generating real revenue. Rebellions is also pursuing a dedicated inference chip path. From SambaNova and Etched in the U.S., to DeepSeek's custom silicon and Hygon's full-stack in China, to Rebellions in Korea, the global AI inference chip landscape is shifting from a NVIDIA monopoly to a diverse, multi-region, multi-architecture, multi-scenario playing field.
The Inference-Era Compute Landscape: From "One Superpower" to "Heterogeneous Warring States"
We are witnessing a fundamental restructuring of AI compute infrastructure.
The training era (2020–2025) had a simple logic: Scaling Laws dictated that whoever could buy the most H100s/H200s/B100s could train the strongest models. NVIDIA was the undisputed king of this era — CUDA ecosystem moats, TSMC advanced process priority, NVLink high-speed interconnects left competitors with virtually no opening. AMD's MI300 series had competitive performance but a massive ecosystem gap; Google TPUs were for internal use only; most AI chip startups remained at the PowerPoint stage.
But the inference era plays by completely different rules. First, inference doesn't need the most advanced process nodes. Training demands extreme compute density and memory bandwidth; inference has far lower process requirements, and mature nodes can deliver competitive performance-per-watt. This opens the door for more players. Second, inference scenarios are highly fragmented. Cloud inference, private data center inference, edge inference, on-device inference — each scenario has different latency, power, and cost requirements, and no single chip can serve all scenarios well. Third, software stack barriers are falling. Inference frameworks like PyTorch, ONNX Runtime, vLLM, and TensorRT-LLM are maturing rapidly, drastically reducing software adaptation costs for new chips.
SambaNova's SN50 choosing deep collaboration with Intel Xeon on a heterogeneous approach, rather than building a standalone chip from scratch, embodies this trend. Liang repeatedly emphasized the Intel partnership in his interview — leveraging Intel's manufacturing scale, supply chain, and customer channels means SambaNova doesn't need to build an entire system from zero. Intel also needs SambaNova — Xeon CPUs need dedicated accelerator chips for AI inference to compete effectively against NVIDIA Grace Hopper.
Microsoft was also recently reported to be joining the AI cost-cutting trend by relying more on its own models. TechCrunch reported that Microsoft is reducing the proportion of OpenAI model calls it uses, shifting toward its in-house Phi series and MaaS models. This sends an equally clear signal: even OpenAI's largest investor is diversifying for inference costs — nobody wants to be locked into a single vendor.
Endgame: A Three-Layer Cake Is Forming for Inference Infrastructure
Pulling these threads together, the endgame outline for AI inference infrastructure is already clear: a three-layer cake.
Layer one: hyperscale public cloud inference. AWS, Azure, GCP, Alibaba Cloud, Tencent Cloud will continue providing general-purpose inference serving SMBs, elastic workloads, and non-sensitive scenarios. The core competitive advantages here are scale and ease of use. NVIDIA GPUs will remain dominant in this layer but will see share fall from near-monopoly to 50–60%.
Layer two: industry/sovereign private inference infrastructure. Banks, governments, healthcare, and large manufacturers will increasingly build their own private AI inference clusters. Dedicated inference chip companies like SambaNova, Etched, and Rebellions, along with domestic compute vendors like Hygon and Huawei Ascend, will compete head-to-head with NVIDIA in this layer. Core competitive advantages here are security and control, TCO, and scenario-specific performance optimization. JPMorgan choosing SambaNova is the starting gun for this layer's explosion.
Layer three: on-device and edge inference. Phones, PCs, cars, industrial equipment, robots — these scenarios require ultra-low latency, extreme privacy, and offline availability. Models will increasingly run on-device rather than in the cloud. Qualcomm, MediaTek, Apple, Hygon, and Horizon Robotics will compete fiercely in this layer.
For AI industry practitioners and investors, this structural shift sends several clear signals. First, inference chips are a bigger market than training chips — training tops out at tens of billions annually; inference is a hundreds-of-billions, even trillion-dollar market that lasts far longer. Second, "full self-development" isn't the only path — SambaNova's Intel partnership proves heterogeneous computing and strong alliances may be more efficient than vertical integration. Third, the window is opening for Chinese inference chip vendors — DeepSeek's custom silicon and Hygon's full-stack push aren't about beating NVIDIA at training; they're about capturing domestic market share in the much larger inference market.
The $11 billion valuation is just the beginning. The independence war for AI inference infrastructure has just fired its first shot.
See you tomorrow.
SambaNova · 110亿估值 · SN50 · 摩根大通 · AI推理芯片 · DeepSeek自研芯片 · 海光端侧AI · 私有部署 · 异构计算 · 算力格局重构 · Intel合作 · Rebellions IPO
SambaNova · $11B valuation · SN50 · JPMorgan Chase · AI inference chips · DeepSeek custom chips · Hygon edge AI · private deployment · heterogeneous computing · compute restructuring · Intel partnership · Rebellions IPO