没有发布会,没有PPT,没有预热海报。
8月13日凌晨,DeepSeek的API页面悄悄更新了一个版本号:DeepSeek-V4-Pro-0813。模型名没变,调用方式没变,文档没变——但你今天调的API,已经不是昨天那个模型了。
最直接的证据是一个数字的跳跃:DeepSWE跑分从12.8直接跳到62.7,翻了将近5倍。一个月前的预览版在正式版面前,像是另一个物种。
这不是一次常规的小版本迭代。这是大模型行业一个标志性的转向——过去两年卷的是推理能力:奥数题、编程竞赛、知识问答。而这次暴涨的每一项跑分,都指向同一个方向:让模型在真实环境里,长时间、多步骤地完成任务。
用行话说,这叫Agent能力。用大白话说:以前的模型是学霸,这次升级的方向是长工。
翻了5倍的,不是智商是干活能力
先看一组对比。左边是一个月前的预览版(Preview),右边是8月13日上线的正式版:
Terminal Bench从72.1到87.9,提升15.8分;DeepSWE从12.8到62.7,翻近5倍;AutomationBench从12.8到31.8,提升19分;DSBench-FullStack从41.8到71.1,提升29.3分;DSBench-Hard从31.1到67.2,提升36.1分;Cybergym从52.7到83.3,提升30.6分。
这张表里最值得多看一眼的,不是涨幅,而是涨幅出现的位置。
Terminal Bench考的是模型在终端里真实操作的能力,DeepSWE考的是软件工程任务,AutomationBench考的是自动化流程,Cybergym考的是安全攻防——这些全是"干活"的测试,不是"答题"的测试。
"过去两年大模型卷的是'能不能答对',现在开始卷的是'能不能把活干完干好'。这不是同一场考试。"—— 一位AI行业从业者的判断
为什么这个转向重要?因为推理能力再强,如果不能在真实环境中持续执行任务,它的商业价值就始终停留在"聊天助手"的层面。而Agent能力——也就是模型能调用工具、操作电脑、完成多步骤任务的能力——才是大模型真正从"玩具"变成"生产力工具"的关键一跃。
DeepSeek这次深夜升级释放的信号非常明确:它不再满足于做"便宜的推理模型",它要直接切入Agent主战场。
横向对比:无短板才是最难做到的
自家预览版被吊打,只能说明进步快。真正的问题是:放到全球牌桌上,DeepSeek V4 Pro站在什么位置?
把正式版和目前第一梯队的几个模型摆在一起看:Terminal Bench上,Kimi-K3以88.3分占据榜首,V4 Pro是87.9,差0.4分——这个差距比一次测试的随机波动大不了多少。HLE(带工具)上,Fable 5的63.0仍然领先。NL2Repo上,Opus-4.8的69.7保持优势。Cybergym 83.3分,则是目前已公布的最高水平之一。
这张表怎么读?看两点。
第一,单项未必全是第一。每一个细分项目都有更强的选手——有的编程强,有的推理强,有的安全攻防强。V4 Pro在任何一个单项上都不是绝对的王者。
第二,也是更关键的一点:没有任何一项掉出第一梯队。
这恰恰是Agent模型最难做到的事。推理模型可以偏科——数学特别强、写作一般般,照样有卖点。但Agent要在真实环境里连续干活:写代码、调工具、查资料、改bug、跑流程,任何一个环节掉链子,整个任务就崩了。
"单项未必全是第一,但没有任何一项掉出第一梯队——这才是Agent模型最难做到的无短板。"—— 一位AI技术博主的评论
用一句话概括现在的格局:国产AI离海外旗舰,只差一步。而且这一步,是以"月"为单位在缩短的。
如果说半年前,国产大模型还在"追赶"海外旗舰,那么现在的格局已经变成了"交替领先"——你在这个基准上领先,我在那个基准上领先。第一梯队的门槛,已经不是美国公司专属了。
地板价的算盘:1M上下文的生意经
说完能力,说钱。这次价格牌,打得很有讲究。
先看给了什么:100万token上下文,38.4万token最大输出。什么概念?一百万token大约能塞进去七八本长篇小说,或者一个中型项目的全部代码。对于要长时间跑任务的Agent来说,长上下文和大输出不是锦上添花,是硬刚需——上下文不够长,干到一半就把前面的指令忘了。
再看收多少。V4 Pro定价:缓存命中输入0.025元/百万token,未命中输入3元/百万token,输出6元/百万token,并发上限500。
两个细节值得注意。
一是缓存命中价压到了0.025元。重复上下文走缓存,成本几乎可以忽略——这是明摆着在鼓励开发者把长任务、多轮对话的场景往上搬。对比海外同级旗舰动辄几十美元的定价,这个价格基本就是零头。同样的活,只要别人一个零头。
二是并发上限从Flash的2500收到了500。为什么?能力越强,单次推理烧的算力越贵,尤其是默认开启的思考模式。从并发上限的收缩也能侧面看出来:V4 Pro的算力成本比V4 Flash高了不少,但DeepSeek仍然把价格压在了一个极具竞争力的水平。
这背后的商业逻辑很清晰:用低价换量,用量练模型,用模型数据再练出更强的能力。DeepSeek不是慈善机构,它的低价策略是一种投资——用今天的低价,换明天的市场份额和数据积累。
模型战争的下半场:从比聪明到比干活
DeepSeek V4 Pro的这次深夜升级,放在更宏观的行业背景下看,意义远不止一个模型的更新。
它标志着大模型行业的竞争,正在发生一次根本性的转向。
上半场的主题是"谁更聪明"。比参数规模、比推理能力、比基准测试跑分。各家公司的发布会,核心叙事都是"我们的模型在某某榜单上超越了某某"。用户关心的也是"这个模型做题够不够厉害"。
下半场的主题正在变成"谁更能干活"。比Agent能力、比工具调用、比真实场景下的任务完成率。基准测试不再是MMLU和GSM8K,而是Terminal Bench、DeepSWE、Cybergym这些真正考验"做事能力"的测试集。
这个转向对行业意味着什么?
第一,评价体系的切换。过去我们用"聪不聪明"来评价一个模型,未来我们会用"能不能干活"来评价。这两个维度有关联,但不完全重合——一个模型可能推理题做得很好,但一到真实环境就各种翻车。
第二,商业模式的切换。纯聊天模型的商业化天花板已经逐渐清晰——订阅制、API调用,这些模式的增长曲线正在放缓。而Agent模型直接对接的是真实工作流——编程、办公、运营、客服——每一个都是万亿美元级别的市场。
第三,竞争格局的重洗。上半场赢了的公司,下半场未必还能赢。推理能力强不等于Agent能力强,数据积累多不等于工具链完善。Agent是一个系统工程,不是光靠模型就能搞定的。
"当模型名没变、里子却换了一代的时候,你就知道这个行业的迭代速度已经快到了什么程度。去年的最强模型,今年可能连第二梯队都进不去。"—— 一位AI投资人的观察
DeepSeek这次选择在深夜静默更新,不发公告、不做发布会,本身就是一种自信的表现——产品够硬的时候,营销是最不重要的那一环。一张跑分图,胜过十场发布会。
但竞争也远没有结束。V4 Pro的Agent能力上来了,那V5呢?OpenAI的下一个版本呢?Anthropic的Claude Next呢?这场从"比聪明"到"比干活"的战争,才刚刚进入白热化阶段。
可以确定的是:未来几个月,我们会看到越来越多的模型在Agent基准上你追我赶。对于用户来说,这是最好的时代——模型越来越强,价格越来越低,选择越来越多。
对于大模型公司来说,这也是最坏的时代——你刚觉得自己领先了,转头就被别人深夜换芯超过去。
明天见。
No launch event. No slides. No teaser posters.
In the early hours of August 13, DeepSeek's API page quietly updated a version number: DeepSeek-V4-Pro-0813. The model name stayed the same, the API call stayed the same, the docs stayed the same - but the API you're calling today isn't the same model as yesterday.
The most direct evidence is a single number's leap: DeepSWE score jumped from 12.8 straight to 62.7 - nearly 5x. The preview version from a month ago, next to the official release, looks like a different species entirely.
This isn't a routine minor version bump. This is a landmark pivot for the entire LLM industry. For the past two years, the race was about reasoning - Olympiad math, programming contests, knowledge QA. But every benchmark that skyrocketed this time points in the same direction: making models work for extended periods, across multiple steps, in real environments.
In industry jargon: Agent capability. In plain English: previous models were honor students; this upgrade's direction is blue-collar workers.
What Jumped 5x Wasn't IQ - It Was Getting Work Done
Let's compare. On the left, the preview version from a month ago. On the right, the official release that went live August 13:
Terminal Bench: 72.1 to 87.9, +15.8. DeepSWE: 12.8 to 62.7, nearly 5x. AutomationBench: 12.8 to 31.8, +19.0. DSBench-FullStack: 41.8 to 71.1, +29.3. DSBench-Hard: 31.1 to 67.2, +36.1. Cybergym: 52.7 to 83.3, +30.6.
What's most notable in this table isn't the gains themselves - it's where the gains are happening.
Terminal Bench tests a model's ability to actually operate inside a terminal. DeepSWE tests software engineering tasks. AutomationBench tests automated workflows. Cybergym tests security offense and defense. All of these are "getting work done" tests, not "answering questions" tests.
"For two years LLMs competed on 'can you get the right answer.' Now they're competing on 'can you get the job done.' These aren't the same exam." - An AI Industry Practitioner
Why does this pivot matter? Because no matter how strong reasoning capability is, if a model can't continuously execute tasks in real environments, its commercial value remains stuck at the "chat assistant" level. Agent capability - a model's ability to call tools, operate a computer, complete multi-step tasks - is the critical leap from "toy" to "productivity tool."
The signal DeepSeek is sending with this midnight upgrade is clear: it's no longer satisfied being a "cheap reasoning model." It's going straight for the Agent main battlefield.
Side by Side: No Weaknesses Is the Hardest Feat
Beating your own preview version only proves you've improved fast. The real question: where does DeepSeek V4 Pro stand at the global table?
Pitting the official release against the current first-tier models: on Terminal Bench, Kimi-K3 leads at 88.3; V4 Pro sits at 87.9 - a 0.4 difference, barely more than test noise. On HLE with tools, Fable 5's 63.0 still leads. On NL2Repo, Opus-4.8's 69.7 holds the edge. On Cybergym, 83.3 ranks among the highest published scores.
How to read this chart? Two observations.
First, it's not first in every category. Every sub-discipline has a stronger player - some excel at coding, some at reasoning, some at security. V4 Pro isn't the undisputed king of any single benchmark.
Second - and this is the more crucial point - it doesn't fall out of the first tier in any of them.
This is precisely the hardest thing for an Agent model to achieve. Reasoning models can afford to be lopsided - great at math, mediocre at writing, and still sell. But an agent working through real tasks has to code, use tools, look things up, fix bugs, run pipelines. Fail at any one link and the whole task breaks.
"Not first in everything, but never outside the top tier - that's the hardest no-weakness bar for Agent models." - An AI Tech Blogger
Summarizing the current landscape in one sentence: Chinese AI is one step away from overseas flagships. And that step is closing, measured in months.
If half a year ago Chinese LLMs were still "catching up" to overseas flagships, the landscape has now shifted to "alternating leads." You lead on this benchmark, I lead on that one. The first tier is no longer an exclusive club for American companies.
Floor-Price Strategy: The Economics of 1M Context
Capability covered, now let's talk money. This pricing play is calculated.
First, what you get: 1 million token context window, 384,000 token max output. To put that in perspective: a million tokens is roughly seven or eight full-length novels, or the entire codebase of a medium-sized project. For agents running long tasks, long context and large output aren't nice-to-haves - they're hard requirements. Too short a context, and the agent forgets earlier instructions halfway through.
Now what you pay. V4 Pro pricing: cached input 0.025 yuan/M tokens, uncached input 3 yuan/M tokens, output 6 yuan/M tokens, concurrency limit 500.
Two details stand out.
First, cached input priced at 0.025 yuan. Repeated context goes through cache and costs essentially nothing - this is clearly encouraging developers to move long tasks and multi-turn conversations onto this model. Compared to overseas flagship prices in the tens of dollars, this price is literally a fraction. The same work, for pocket change.
Second, concurrency dropped from 2500 (Flash) to 500. Why? The more capable the model, the more compute each inference burns - especially with thinking mode enabled by default. The concurrency contraction also tells you something: V4 Pro's compute cost is significantly higher than V4 Flash's, yet DeepSeek still compressed the price to a brutally competitive level.
The business logic is clear: low prices for volume, volume for data, data for even better capability. DeepSeek isn't a charity - its low-price strategy is an investment. Today's low price buys tomorrow's market share and data accumulation.
The Model War's Second Half: From Smarts to Work
DeepSeek V4 Pro's midnight upgrade, viewed against the broader industry backdrop, means far more than a single model update.
It signals that competition in the LLM industry is undergoing a fundamental pivot.
The first half's theme was "who's smarter". Competing on parameter scale, reasoning capability, benchmark scores. Every company's keynote centered on "our model surpassed X on benchmark Y." Users cared about "can this model ace the test."
The second half's theme is becoming "who gets work done". Competing on Agent capability, tool use, real-world task completion. Benchmarks aren't MMLU and GSM8K anymore - they're Terminal Bench, DeepSWE, Cybergym, tests that actually measure "getting things done."
What does this pivot mean for the industry?
First, a shift in evaluation. We used to judge models by "how smart they are." Going forward, we'll judge by "how much work they can do." These dimensions correlate but aren't identical - a model can nail reasoning questions yet completely fall apart in real environments.
Second, a shift in business models. The commercial ceiling of pure chat models is becoming clearer - subscriptions, API calls, these growth curves are flattening. Agent models directly plug into real workflows - coding, office operations, customer service - each a multi-trillion-dollar market.
Third, a reshuffle of the competitive landscape. Companies that won the first half won't necessarily win the second. Strong reasoning doesn't equal strong Agent capability. Lots of data doesn't equal a complete toolchain. Agent systems are systems engineering - you can't just throw a model at it and call it done.
"When the model name stays the same but the inside is a whole new generation, you know how fast this industry iterates. Last year's strongest model might not even make second tier this year." - An AI Investor
DeepSeek choosing to release this silently in the middle of the night - no announcement, no press event - is itself a display of confidence. When the product speaks for itself, marketing is the least important thing. One benchmark chart beats ten launch events.
But the competition is far from over. V4 Pro's Agent capability is up - but what about V5? What about OpenAI's next version? What about Anthropic's Claude Next? This war, shifting from "who's smarter" to "who gets work done," is just entering its white-hot phase.
One thing is certain: in the months ahead, we'll see more and more models trading leads on Agent benchmarks. For users, this is the best of times - models keep getting stronger, prices keep dropping, options keep multiplying.
For LLM companies, it's also the worst of times - just when you think you're ahead, someone swaps cores on you at midnight and surges past.
See you tomorrow.