大模型

大模型定价
两极分化

LLM Pricing
Goes Bipolar

DeepSeek V4-Pro杀到GPT-5.5的1/35,Claude Fable 5翻倍涨价——大模型市场正在残酷分层:要么极致便宜,要么极致聪明,中间的没有活路。

DeepSeek V4-Pro slashes to 1/35th of GPT-5.5 pricing; Claude Fable 5 doubles prices -- the LLM market is brutally stratifying: be ultra-cheap or ultra-smart; the middle has no path to survival.

No.003 2026.06.26 约 5 分钟阅读 ~5 min read

大模型API的价格战打到了一个荒诞的地步。

DeepSeek V4-Pro的缓存命中输入价格降到了0.025元/百万tokens——作为对比,GPT-5.5的同等输入价格约为0.88元。也就是说,DeepSeek的价格已经杀到了GPT的1/35。另一边,Anthropic的Claude Fable 5不降价反涨价,输入$10/输出$50,价格直接翻倍,走的是"最贵但最好"的高端路线。

一个往地板上砸,一个往天花板上冲。中间地带的模型厂商,正在被两边挤压。

两种活法

Altimeter Capital的合伙人给出了一个直白的判断:大模型市场最终只剩两种活法——"要么足够出色,要么非常便宜"。

足够出色的,比如Claude和GPT的旗舰版本,它们靠最顶尖的推理能力、最强的代码生成、最可靠的长文本理解,赚取头部用户的高溢价。数据很说明问题:Claude Code头部10%的超级用户贡献了80-90%的营收。这些用户不在乎单价是$5还是$50,他们在乎的是"能不能把活干对"。

非常便宜的,比如DeepSeek和各种开源微调模型,它们靠极致的成本控制承载海量token消耗。80%的日常任务——文本分类、简单摘要、常规客服——不需要旗舰模型的智商,一个够快够便宜的模型就足够了。

"前沿模型攫取90%经济价值,开源廉价模型承载80% token消耗。中间的,先死。" —— 行业终局预判

智能路由成为核心形态

这种分化催生了一个新的关键层:智能路由(Model Router)。

法律AI独角兽Harvey已经给出了一个范本:他们通过混合路由方案——简单任务分配给开源微调模型,复杂任务路由到Opus 4.7/4.8——在部分场景上超越了纯Opus的效果,成本却大幅降低。这不是"一个模型打天下"的思路,而是"让合适的模型做合适的事"。

未来的AI应用架构中,模型选择本身将成为一个核心能力。你不需要绑定某一家模型厂商,你需要一个聪明的"调度员"——它知道什么任务该用什么模型,知道什么时候该省钱、什么时候该砸钱。

对开发者意味着什么

对于普通开发者和创业者来说,定价分化带来了两个明确的启示:

第一,不要在"做一个比GPT便宜一点但差不多好"的模型上浪费时间。这个位置已经不存在了。要么你能做到DeepSeek那样的极致成本,要么你能做到Claude那样的极致能力,否则就是死路一条。

第二,智能路由层是一个被低估的机会。当模型成为像水电一样的公用事业,"帮你选对模型"这件事本身就有了商业价值。就像云计算时代诞生了CloudHealth这样的成本优化公司,AI时代也会诞生模型路由和成本优化的新玩家。

价格战还会继续打下去。但结局已经隐约可见:大模型不再是一个"赢家通吃"的市场,而是一个分层的市场——顶层卖智商,底层卖算力,中间层卖路由和服务。


明天见。

LLM API price wars have reached an absurd point.

DeepSeek V4-Pro's cache-hit input price dropped to 0.025 RMB per million tokens -- by comparison, GPT-5.5's equivalent input price is roughly 0.88 RMB. That means DeepSeek's price has slashed to 1/35th of GPT. On the other side, Anthropic's Claude Fable 5 didn't cut prices but raised them -- input $10/output $50, prices effectively doubled, pursuing a "most expensive but best" premium route.

One slams into the floor; the other shoots for the ceiling. Model vendors in the middle are getting squeezed from both sides.

Two Ways to Survive

An Altimeter Capital partner put it bluntly: the LLM market ultimately leaves only two survival modes -- "either be outstanding, or be very cheap."

The outstanding ones, like flagship Claude and GPT, earn premium pricing from top users through cutting-edge reasoning, strongest code generation, most reliable long-context understanding. The data speaks volumes: Claude Code's top 10% power users contribute 80-90% of revenue. These users don't care if the unit price is $5 or $50; they care about "getting the job done right."

The very cheap ones, like DeepSeek and various open-source fine-tuned models, carry massive token consumption through extreme cost control. 80% of everyday tasks -- text classification, simple summarization, routine customer service -- don't need flagship IQ; a fast, cheap model is enough.

"Frontier models capture 90% of economic value; cheap open-source models carry 80% of token consumption. The middle? Dies first." -- Industry endgame prediction

Model Routing Becomes the Core Pattern

This stratification is spawning a critical new layer: the Model Router.

Legal AI unicorn Harvey has already provided a template: through a hybrid routing scheme -- assigning simple tasks to open-source fine-tunes and routing complex tasks to Opus 4.7/4.8 -- they exceeded pure Opus results on some scenarios while slashing costs. This isn't a "one model rules all" approach; it's "let the right model do the right job."

In future AI application architectures, model selection itself becomes a core competency. You don't need to bind to one model vendor; you need a smart "dispatcher" -- it knows what model to use for what task, knows when to save money and when to spend it.

What It Means for Developers

For regular developers and founders, pricing polarization brings two clear takeaways:

First, don't waste time building a model that's "a bit cheaper than GPT but about as good." That position no longer exists. Either you can achieve DeepSeek-level extreme cost, or Claude-level extreme capability -- otherwise it's a dead end.

Second, the model routing layer is an underappreciated opportunity. When models become utilities like water and electricity, "helping you pick the right model" itself becomes commercially valuable. Just as the cloud computing era spawned cost-optimization companies like CloudHealth, the AI era will spawn new players in model routing and cost optimization.

Price wars will continue. But the ending is coming into view: LLMs won't be a "winner takes all" market, but a stratified one -- the top tier sells IQ, the bottom tier sells compute, and the middle sells routing and services.


See you tomorrow.

前沿模型攫取90%经济价值,开源廉价模型承载80% token消耗。中间的,先死。

—— 行业终局预判
大模型定价分层趋势 · 智能路由架构
LLM pricing stratification trends, model router architecture
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 5 个源信号生成,经编辑部人工审核。素材来源包括:36氪大模型定价分析、Altimeter投资研判、Harvey混合路由方案。

Generated by Dawn Vision's cognitive engine from 5 source signals, editorially reviewed. Sources include: 36Kr LLM pricing analysis, Altimeter investment thesis, Harvey hybrid routing approach.