算力的价格,涨得比油价还猛。
SemiAnalysis的最新报告显示,英伟达H100 GPU的一年期租赁价格,从2025年10月的每GPU每小时约1.70美元,涨到了2026年3月的2.35美元,涨幅接近40%。更夸张的是——现有所有类型GPU的按需租赁产能,几乎全部售罄。
拿旺季抢火车票来比喻都不够贴切了——2026年初,获取GPU算力的难度,可能比春运抢票还高。
为什么越涨越缺货?
正常的市场逻辑是:价格涨了,需求就会降,供需重新平衡。但GPU市场好像不太吃这一套。
原因在于需求的刚性。Anthropic的Claude系列模型年化营收从90亿美元涨到250亿美元以上,OpenAI的规模更大。字节跳动、谷歌的媒体生成类AI工具快速普及,带动了对算力的激增需求。这些都是已经跑通商业模式的公司——对它们来说,GPU不是成本项,是生产资料,是收入的前提。
更关键的变量是多智能体工作流。Agent的大规模应用,带来了token消耗的指数级增长。以前一个用户问一个问题,消耗几十个token。现在一个Agent任务可能要跑几十步推理、调用多个工具、生成大量中间结果——token消耗量直接翻了几个数量级。
超高的投资回报率,使得算力需求趋于刚性。价格涨了?那就少赚一点,但不能不用。这种刚需属性,让GPU租赁市场陷入了一个自我强化的循环:越缺货,价格越涨;价格越涨,云厂商越愿意提前锁定更多产能;产能被锁得越多,现货就越缺货。
老款不跌反涨的反常信号
还有一个值得注意的反常现象:新芯片发布,老款不但没降价,反而也在涨。
市场原本预期更高能效的Blackwell系列发布后,H100等旧款芯片的价格会下跌。结果呢?老款GPU需求依旧强劲,现货价格没有任何松动迹象。甚至很多两三年前的H100租赁合约,都在以原价续签,有些还续到了2028年。
为什么?因为需求增长的速度,远快于新芯片产能释放的速度。Blackwell的交付周期已经延长到2026年下半年,甚至8-9月的产能都被预订完了。在新产能跟不上的情况下,老款芯片的价值不仅没有贬值,反而因为立刻能用而变得更稀缺。
上游产业链也在推波助澜。DRAM和NAND闪存价格在2026年一季度跳涨,核心组件涨价推高了AI服务器的整体成本。服务器OEM厂商大幅上调报价,涨幅甚至超过了组件成本的上涨幅度,直接压缩了算力集群项目的预期收益,让新增供给更加犹豫。
这种情况下,算力荒短期内看不到缓解的迹象。真正的拐点,可能要等到Blackwell大规模落地、或者推理效率的提升能实质性抵消需求增长的时候。
在此之前,有卡的,就是大哥。
明天见。
Compute prices are rising faster than oil.
SemiAnalysis' latest report shows that one-year lease prices for NVIDIA H100 GPUs climbed from roughly $1.70 per GPU-hour in October 2025 to $2.35 by March 2026 — nearly a 40% increase. More strikingly: on-demand rental capacity for essentially all GPU types is sold out.
Comparing it to 'harder to get than train tickets during Spring Festival' doesn't even do it justice. In early 2026, getting GPU compute is probably harder than securing festival transport in China.
Why Do Prices Keep Rising Instead of Cooling Demand?
Normal market logic says: as prices go up, demand goes down, and supply and demand rebalance. The GPU market seems to play by different rules.
The reason is demand inelasticity. Anthropic's Claude series went from $9B to over $25B in annualized revenue. OpenAI is even bigger. ByteDance and Google are seeing rapid adoption of generative media tools, driving surging compute demand. These are companies with proven business models — for them, GPUs aren't a cost line item. They're the means of production. The prerequisite for revenue.
The more critical variable is multi-agent workflows. Massive adoption of agents has brought exponential growth in token consumption. Before, a user asked one question and burned a few dozen tokens. Now one agent task might run dozens of reasoning steps, call multiple tools, and generate huge amounts of intermediate output — total token consumption multiplies several times over.
With ROI this high, compute demand becomes inelastic. Prices go up? Fine — just make less margin. But you can't not have it. This rigidity has trapped the GPU rental market in a self-reinforcing cycle: scarcer supply → higher prices → cloud providers lock in more capacity ahead of time → even less spot availability → even more scarcity.
The Strange Signal: Older GPUs Are Going Up, Not Down
There's another反常 signal worth noting: when new chips launch, older models usually drop in price. Not this time. They're going up.
The market had expected that higher-efficiency Blackwell GPUs would push H100 and other older-generation prices down. Instead, demand for legacy GPUs has stayed strong, and spot prices show no signs of softening. Even H100 leases from two or three years ago are being renewed at original prices — sometimes extending all the way to 2028.
Why? Because demand is growing faster than new chip capacity can be released. Blackwell delivery lead times have stretched into the second half of 2026, and even August-September allocation is already spoken for. When new supply can't keep up, the value of immediately available older chips doesn't depreciate — it becomes more scarce.
The upstream supply chain is amplifying the effect. DRAM and NAND flash prices jumped in Q1 2026. Rising component costs push up overall AI server prices. OEMs have raised quotes significantly — in some cases more than the component increase itself — directly squeezing expected returns on new compute cluster projects and making new supply additions even more hesitant.
Under these conditions, there's no relief from the compute shortage in sight. The real turning point might not come until Blackwell scales up massively, or until inference efficiency gains substantially offset demand growth.
Until then — whoever has the cards, calls the shots.
See you tomorrow.