9月10日,Cognition发布了SWE-2编程模型。这不是一次常规的模型更新——它标志着AI编程模型的价格战正式开打。
核心数据:SWE-2在Cognition自建的FrontierCode 1.1 Main基准上得分50.0%,仅比Anthropic的Fable 5.1低0.9个百分点(50.9%),比GPT-6 Astra低3.3个百分点(53.3%)。但SWE-2的推理成本比Fable 5.1低64%,是GPT-6 Astra的四分之一。
Kimi K3做基座:开源模型的逆袭
SWE-2最引人注目的细节是它的基座模型:Moonshot的Kimi K3。
这意味着Cognition没有选择GPT-6 Astra或Claude Fable 5.1作为基座,而是选择了一个开源模型。然后通过RL(强化学习)后训练,将Kimi K3的编程能力提升到了接近前沿闭源模型的水平。
这个选择有两层含义:第一,开源模型在编程领域已经具备了足够的基础能力,后训练可以进一步放大这个优势;第二,使用开源模型可以大幅降低成本,这正是SWE-2能够便宜64%的核心原因。
Terminal-Bench 2.1:SWE-2的真正主场
虽然在FrontierCode上略逊于Fable 5.1和Astra,但SWE-2在另一个基准上表现出色:Terminal-Bench 2.1得分92.8分,领先所有竞品。
Terminal-Bench是一个更贴近真实编程场景的基准,测试的是模型在终端环境中完成复杂任务的能力。SWE-2在这个基准上的领先,说明它在实际编程工作流中可能比基准分数显示的更有竞争力。
价格战的信号
SWE-2的发布释放了一个明确的信号:AI编程模型的定价正在被重新定义。
过去一年,AI编程模型的价格一直在下降,但幅度有限。SWE-2用开源基座+RL后训练的方式,实现了接近前沿模型的性能,同时大幅降低成本。这个模式如果被验证可行,将对整个AI编程市场产生深远影响。
对于开发者来说,这是好消息。当编程AI的成本降低64%时,更多开发者可以负担得起AI辅助编程,AI编程工具的普及速度将大幅加快。
"SWE-2证明了一件事:开源模型+精细后训练,可以在编程领域挑战闭源前沿模型。"—— 一位AI编程工具开发者
Cognition的战略
SWE-2的发布也体现了Cognition的战略调整。
Cognition的旗舰产品Devin是AI编程Agent,之前一直依赖Claude和GPT系列模型。现在,Cognition开始自研模型,并用开源模型做基座。这意味着Cognition正在从"Agent框架公司"向"模型+Agent全栈公司"转型。
这个转型的核心逻辑是:控制成本,掌控体验。当你的Agent依赖第三方模型时,成本和体验都受制于人。自研模型可以让你根据Agent场景优化模型,同时大幅降低推理成本。
明天见。
On September 10, Cognition launched the SWE-2 coding model. This isn't a routine model update — it marks the official start of the AI coding model price war.
Key data: SWE-2 scored 50.0% on Cognition's own FrontierCode 1.1 Main benchmark, only 0.9 percentage points below Anthropic's Fable 5.1 (50.9%) and 3.3 below GPT-6 Astra (53.3%). But SWE-2's inference cost is 64% lower than Fable 5.1 and one-quarter of GPT-6 Astra's.
Kimi K3 as Base: Open-Source Model's Comeback
SWE-2's most striking detail is its base model: Moonshot's Kimi K3.
This means Cognition didn't choose GPT-6 Astra or Claude Fable 5.1 as the base, but an open-source model. Then through RL (reinforcement learning) post-training, it elevated Kimi K3's coding capabilities to near-frontier closed-source model levels.
This choice has two implications: First, open-source models now have sufficient base capability in coding, and post-training can further amplify this advantage. Second, using open-source models dramatically reduces costs — the core reason SWE-2 is 64% cheaper.
Terminal-Bench 2.1: SWE-2's True Home Turf
While slightly behind Fable 5.1 and Astra on FrontierCode, SWE-2 excels on another benchmark: Terminal-Bench 2.1 at 92.8 points, leading all competitors.
Terminal-Bench is a benchmark closer to real-world coding scenarios, testing a model's ability to complete complex tasks in terminal environments. SWE-2's lead here suggests it may be more competitive in actual coding workflows than benchmark scores indicate.
The Price War Signal
SWE-2's launch sends a clear signal: AI coding model pricing is being redefined.
Over the past year, AI coding model prices have been declining, but modestly. SWE-2 uses the open-source base + RL post-training approach to achieve near-frontier performance while dramatically cutting costs. If this model proves viable, it will have profound implications for the entire AI coding market.
For developers, this is good news. When coding AI costs drop 64%, more developers can afford AI-assisted coding, and AI coding tool adoption will accelerate significantly.
"SWE-2 proves one thing: open-source models + fine-tuned post-training can challenge closed-source frontier models in coding."— An AI coding tool developer
Cognition's Strategy
SWE-2's launch also reflects Cognition's strategic shift.
Cognition's flagship product Devin is an AI coding agent that previously relied on Claude and GPT models. Now, Cognition is developing its own models using open-source bases. This means Cognition is transforming from an "agent framework company" to a "model + agent full-stack company."
The core logic: control costs, own the experience. When your agent depends on third-party models, both cost and experience are at others' mercy. Building your own models lets you optimize for agent scenarios while dramatically reducing inference costs.
See you tomorrow.
SWE-2证明了一件事:开源模型+精细后训练,可以在编程领域挑战闭源前沿模型。
—— 一位AI编程工具开发者
SWE-2 proves one thing: open-source models + fine-tuned post-training can challenge closed-source frontier models in coding.
— An AI coding tool developer
Cognition · SWE-2 · Devin · Kimi K3 · AI编程 · FrontierCode · Terminal-Bench · 价格战 · 开源
Cognition · SWE-2 · Devin · Kimi K3 · AI coding · FrontierCode · Terminal-Bench · price war · open source
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 12 个源信号生成,经编辑部人工审核。素材来源:Cognition官方博客、ai-tldr.dev、runtimewire、OSCHINA。
This article was generated by the Dawn Vision cognitive engine processing 12 source signals, with human editorial review. Sources: Cognition Blog, ai-tldr.dev, runtimewire, OSCHINA.