Focus · 焦点

GPT-5.6三档模型齐发
ultra多Agent模式上线

GPT-5.6 Trio Launches
Ultra Multi-Agent Mode Goes Live

OpenAI发布GPT-5.6家族:Sol旗舰53.6分领先Fable 5达13.1分,Terra/Luna以1/16成本达成同级性能,ultra模式4 Agent并行,同日成为M365 Copilot首选模型。

OpenAI launches GPT-5.6 family: flagship Sol hits 53.6 on Agents' Last Exam, beating Fable 5 by 13.1 points; Terra/Luna match tier performance at 1/16th cost; ultra mode runs 4 agents in parallel — becomes M365 Copilot's preferred model same day.

No.012 2026.07.10 约 10 分钟阅读 ~10 min read

53.6分。这是GPT-5.6 Sol在Agents' Last Exam上拿到的成绩,比Claude Fable 5的adaptive reasoning模式高出13.1分

7月9日,OpenAI正式发布GPT-5.6模型家族——旗舰Sol、均衡Terra、轻量Luna三档齐发,同时推出全新的ultra推理模式,默认调度4个Agent并行处理复杂任务。微软CEO萨蒂亚·纳德拉同日在X平台宣布,GPT-5.6已在Copilot Chat、Cowork、M365应用、GitHub和Foundry全线同步上线,成为Microsoft 365 Copilot的首选模型。

这不是一次普通的模型迭代。发布时间点耐人寻味:OpenAI二号人物Fidji Simo同日宣布因神经免疫系统疾病转任兼职顾问,首席未来学家Joshua Achiam即将月底离职,而OpenAI与微软的"分手传闻"过去一周甚嚣尘上。在高层动荡和关系疑云的关口,OpenAI选择用产品说话。

三档模型+ultra模式:效率是这次的核心命题

GPT-5.6家族最大的变化不是"又强了",而是"同样强但更便宜",或者说"花同样的钱能做更多事"。

旗舰Sol设定了新的智能效率标杆。在Artificial Analysis Coding Agent Index上,Sol max reasoning模式拿到80分的SOTA成绩,比Fable 5高出2.8分,但输出token不到一半,耗时不到一半,成本约为后者的三分之一。在Terminal-Bench 2.1和DeepSWE这两个真实代码库的长程工程测试上,Sol同样刷新纪录。Qodo联合创始人兼CEO Itamar Friedman评价:"在我们的PR基准测试上,Sol用约3倍少的token量击败GPT-5.5,延迟降低约2倍。"

但更具杀伤力的是Terra和Luna。这两款定位中低端的模型,性能表现颠覆了"便宜没好货"的常识:Terra的表现略高于Fable 5,Luna超过Opus 4.8——两者的成本都只有Fable 5的约1/16,耗时约1/3,输出token约一半。Notion联合创始人Simon Last反馈:"许多跑在GPT-5.5上的Agent,换到Terra上表现一样好,成本减半,token少用16%。"

ultra模式是GPT-5.6最具革命性的更新。与以往单模型串行推理不同,ultra默认调度4个Agent并行工作,在BrowseComp、SEC-Bench Pro、Terminal-Bench 2.1三项评测上,4-Agent配置将分数-延迟前沿线大幅推向左上方——更强结果,更短时间。在BrowseComp和SEC-Bench Pro上,16-Agent配置还能进一步提升。这标志着大模型推理正式从"单体智能"进入"多智能体协作"阶段。

OpenAI同时推出了Programmatic Tool Calling(程序化工具调用),模型可以在Responses API中编写和运行轻量程序来协调工具、过滤中间数据、自主决定下一步动作,而不是把每个工具返回结果都传回模型。这大幅减少了token消耗和模型往返次数——Rogo反馈在金融研究任务上输出token减少24%,速度提升28%;PlayCo在Unity场景构建中token减少63.5%,模型轮次减少50.1%。

"GPT-5.6给人的感觉不像是一个聊天助手,更像是一个端到端的技术操作员——检查线上系统、调试问题、修改代码、验证结果、发布产物,在长会话中保持极强的上下文扎根能力。"—— Ian Tracey,Ramp软件工程师

设计判断力跃升:从写代码到做产品

GPT-5.6的另一个关键突破在设计和前端能力

OpenAI在发布中重点展示了Sol的设计判断力:仅凭高层级指令,就能创造出有品味、符合人体工学、功能完整的界面。更强的computer use能力让它不仅生成底层代码或内容,还能检查渲染结果、发现视觉和功能问题、做收尾打磨。从帆船游戏、博物馆网站到室内设计展示,Sol生成的前端界面已经达到了可以直接交付的水准。

Triple Whale CEO AJ Orbach在七项前端基准测试中给了Sol 4.4分(5分制),对比GPT-5.5的4.0分和Claude 4.8的3.5分:"Sol能持续把复杂的电商、仪表盘、产品需求转化为完整的响应式界面,跨桌面和移动端。"Canva AI产品负责人Danny Wu也提到Sol在幻灯片生成上比竞品强约1.6倍token效率。

在知识工作领域,Sol在BrowseComp上达到92.2%的SOTA,OSWorld 2.0达到62.6%——后者超越Opus 4.8的同时输出token少85%。在Microsoft 365场景中,Sol能从零创建完全可编辑的演示文稿,理解参考模板的设计系统(布局、字体、间距、颜色、Slide Master规则),并将这些规范一致地应用到新内容上。ModelML反馈在FinBench的客户工作流中,Sol比Fable少用39%的token,生成的幻灯片更精致、数据可视化更清晰。

这意味着什么?过去AI生成内容最大的问题是"能用但丑"——代码能跑但界面粗糙,文档有内容但排版糟糕,幻灯片逻辑通但没法直接给客户看。GPT-5.6在设计判断力上的跃升,让AI的输出从"半成品"向"可交付成果"迈了一大步。

全线接入微软:分手传闻下的"秀恩爱"

GPT-5.6发布当天最值得玩味的信号,是纳德拉的反应速度。

纳德拉在X上亲自宣布GPT-5.6在Copilot Chat、Cowork、M365、GitHub、Foundry五大产品线同步上线。微软365 Copilot部门EVP Charles Lamanna在官方评价中说:"GPT-5.6标志着Microsoft 365中产物生成能力的进步,生成的输出高度连贯、准确、开箱即用。"

这个时间点和姿态,很难不让人联想到过去一周的"OpenAI与微软分手"传闻。此前有报道称OpenAI在探索脱离Azure独立建设算力基础设施,双方关系出现裂痕。但GPT-5.6的发布和全线接入,至少在产品层面传递了一个明确信号:合作仍在深化

不过OpenAI同时也在释放独立信号。GPT-5.6发布中特别强调了Stripe、Ramp、Shopify、Cisco、Balyasny Asset Management等非微软系企业客户的背书,展示其在企业市场不依附于微软渠道的独立能力。Ramp的Ian Tracey描述了Sol作为"端到端技术操作员"的能力,Shopify的高级应用AI/ML工程师Shane Moran提到Sol在多阶段Codex工作流中对意图的理解优于GPT-5.5。

在定价层面,OpenAI延续了"高效默认"策略。没有公布具体API价格,但官方明确表示Sol在max reasoning下的性价比大幅提升——以Artificial Analysis Intelligence Index为参照,Sol max reasoning在分数接近Fable 5的情况下,耗时少61%,成本约为一半。Terra和Luna的定位更是清晰:让中等和轻量级任务以极低成本获得接近前沿模型的体验。

高层动荡下的产品节奏:OpenAI的"照常营业"

GPT-5.6发布的同一天,两条人事消息为这场产品盛宴投下阴影。

第一条是OpenAI应用业务负责人、内部排名第二的高管Fidji Simo宣布因长期慢性神经免疫系统疾病恶化,离开全职岗位转任兼职顾问。Simo在内部信中表示三个月前开始休病假,经过恢复后决定将更多精力放在健康上。第二条是OpenAI首席未来学家Joshua Achiam宣布将于本月底离职,他在OpenAI效力近九年,是见证公司从非营利安全研究机构转型为商业巨头的元老级人物。

两人的离职原因不同——Simo是健康原因,Achiam的离职被部分观察者视为安全派系持续流失的又一案例——但叠加在一起,加上此前的首席科学家Ilya Sutskever离巢创立SSI、Jan Leike投奔Anthropic,OpenAI的"安全派"人才流失已经成为一个不容忽视的趋势。

但OpenAI选择在同一天发布GPT-5.6这样的重大产品,本身就是一种回应:不管高层怎么变动,产品迭代节奏不会停。这种"show, don't tell"的方式,比任何公关声明都有力。在考虑IPO、追赶Anthropic企业市场、探索算力独立的关键节点,OpenAI需要用持续的产品输出来向投资人、客户和员工证明:机器还在高速运转。

"OpenAI的策略很清晰——当外界在讨论你的组织架构和合作关系时,发布一个让所有竞争对手都要追赶三个月的模型,是最好的回应。"—— Dawn Vision编辑部判断

终局判断:Agentic AI的效率拐点已至

把GPT-5.6家族的发布放在AI产业演进的大背景下看,一个清晰的拐点正在浮现。

2023-2024年是大模型的"能力竞赛"阶段,比拼的是谁的模型更聪明、谁的benchmark分数更高。2025年是"价格战"阶段,DeepSeek把推理成本打到地板价,各家纷纷跟进降价。到了2026年下半年的今天,竞赛已经进入"效率战争"阶段——不是谁更强,而是谁能用更少的token、更低的成本、更短的时间完成同样复杂的任务。

GPT-5.6的三档产品线布局(Sol/Terra/Luna)就是效率战争的产物。这背后是一个朴素的商业逻辑:Agentic工作负载的token消耗量是传统聊天的10-100倍,如果不能把单位token的效率提上去,Agent的大规模商业化就是空中楼阁。ultra模式的多Agent并行、Programmatic Tool Calling的程序化工调用,本质上都是在解决同一个问题:如何让AI Agent在真实生产环境中跑得动、跑得久、跑得便宜

对开发者来说,GPT-5.6释放的信号很明确:第一,不要为每一个任务都调用最强模型,Terra和Luna已经能在大部分场景提供足够好的性能,成本优势巨大;第二,多Agent架构不再是实验性功能,ultra模式的推出意味着OpenAI官方已经把多Agent协作作为前沿推理的默认范式;第三,设计和前端能力是新的竞争高地,AI从"写代码"到"做产品"的跃迁正在发生。

对企业用户来说,GPT-5.6全线接入M365 Copilot意味着AI生产力工具的可用性门槛再次降低。当AI能直接生成符合企业模板规范的幻灯片、报表和文档,而不是产出需要大量人工修改的草稿时,AI在企业的渗透率将迎来真正的非线性增长。

53.6分不是终点。在AI这个指数级迭代的赛道,GPT-5.6的领先窗口期可能只有两到三个月。但OpenAI在高层动荡和关系疑云中按时交出的这份答卷,已经足够说明问题:这家公司的产品引擎,还在全速运转。

明天见。

53.6. That's the score GPT-5.6 Sol posted on Agents' Last Exam — 13.1 points ahead of Claude Fable 5's adaptive reasoning mode.

On July 9, OpenAI officially launched the GPT-5.6 model family — flagship Sol, balanced Terra, and lightweight Luna, all shipping simultaneously — alongside a new ultra reasoning mode that defaults to orchestrating four agents in parallel for complex tasks. Microsoft CEO Satya Nadella took to X the same day to announce that GPT-5.6 had gone live across Copilot Chat, Cowork, M365 apps, GitHub, and Foundry, becoming the preferred model for Microsoft 365 Copilot.

This isn't a routine model iteration. The timing speaks volumes: OpenAI's number-two executive Fidji Simo announced the same day she's moving to an advisory role due to a neuroimmune disease; chief futurist Joshua Achiam is leaving at month's end; and rumors of an OpenAI-Microsoft "split" have swirled for the past week. At the intersection of executive turmoil and relationship uncertainty, OpenAI chose to let the product do the talking.

Three Tiers + Ultra: Efficiency Is the Core Thesis

The biggest shift with GPT-5.6 isn't "it's stronger again" — it's "just as strong, way cheaper," or rather, "same budget, way more work done."

Flagship Sol sets a new intelligence-efficiency benchmark. On the Artificial Analysis Coding Agent Index, Sol in max reasoning mode hit a score of 80, a new SOTA — 2.8 points above Fable 5 — while using fewer than half the output tokens, taking less than half the time, and costing roughly a third less. It also set new SOTA marks on Terminal-Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases. Qodo Co-Founder & CEO Itamar Friedman noted: "On our apples-to-apples PR benchmarks, it beat GPT-5.5 on F1 while using roughly 3x fewer tokens per PR and delivering about 2x lower median latency."

But the real disruption comes from Terra and Luna. These mid- and entry-tier models defy the "you get what you pay for" axiom: Terra performs slightly above Fable 5, and Luna outperforms Opus 4.8 — both at roughly one-sixteenth the cost, in about one-third the time, with about half as many output tokens. As Notion Co-Founder Simon Last put it: "Many agents running GPT-5.5 perform just as well on Terra for half the cost and 16% fewer tokens."

Ultra mode is GPT-5.6's most revolutionary update. Unlike prior serial single-model reasoning, ultra defaults to orchestrating four agents in parallel. Across BrowseComp, SEC-Bench Pro, and Terminal-Bench 2.1, the four-agent configuration pushes the score-latency frontier dramatically up and left — stronger results in less time. On BrowseComp and SEC-Bench Pro, a 16-agent configuration pushes performance even further. This marks the formal transition of large-model reasoning from "solitary intelligence" to "multi-agent collaboration."

OpenAI also introduced Programmatic Tool Calling, which lets models write and run lightweight programs within the Responses API to coordinate tools, filter intermediate data, and decide next steps autonomously — rather than passing every tool response back through the model. This slashes token consumption and round trips: Rogo reported 24% fewer output tokens and 28% faster completion on financial research tasks; PlayCo saw 63.5% fewer total tokens and 50.1% fewer model turns on Unity scene construction.

"GPT-5.6 felt less like a chat assistant and more like an end-to-end technical operator. It could inspect live systems, debug issues, make code changes, validate results, publish artifacts, and carry context across long sessions with strong grounding."— Ian Tracey, Software Engineer, Ramp

Design Judgment Leap: From Writing Code to Building Products

Another critical breakthrough in GPT-5.6 is design and frontend capability.

OpenAI's release heavily showcased Sol's design judgment: from high-level direction alone, it creates tasteful, ergonomic, fully functional interfaces. Stronger computer-use capabilities let it inspect rendered results — not just generate underlying code — catch visual and functional issues, and apply finishing touches. From a sailing game and museum website to interior design presentations, Sol's frontend output has reached shippable quality.

Triple Whale CEO AJ Orbach gave Sol a 4.4 out of 5 across a seven-task frontend benchmark, versus 4.0 for GPT-5.5 and 3.5 for Claude 4.8: "It consistently turned complex ecommerce, dashboard, and product briefs into complete, responsive interfaces across desktop and mobile." Canva Head of AI Products Danny Wu noted Sol was roughly 1.6x more token-efficient at slide generation than competitors.

In knowledge work, Sol scored 92.2% SOTA on BrowseComp and 62.6% on OSWorld 2.0 — surpassing Opus 4.8 on the latter while using 85% fewer output tokens. In Microsoft 365 scenarios, Sol can create fully editable presentations from scratch, infer a reference deck's design system (layouts, typography, spacing, colors, Slide Master rules), and apply those conventions consistently to new material. ModelML reported Sol used 39% fewer tokens per deck than Fable across client workflows, producing more polished slides with clearer data visualizations.

What does this mean? The biggest problem with AI-generated content has been "functional but ugly" — code that runs but looks rough, documents with content but poor formatting, slides that make sense but can't be shown to clients. GPT-5.6's leap in design judgment pushes AI output from "rough draft" toward "shippable deliverable."

Full Microsoft Integration: A Show of Unity Amid Split Rumors

The most telling signal on GPT-5.6 launch day was the speed of Nadella's response.

Nadella personally took to X to announce GPT-5.6 going live simultaneously across five product lines: Copilot Chat, Cowork, M365 apps, GitHub, and Foundry. Microsoft 365 Copilot EVP Charles Lamanna stated officially: "GPT-5.6 marks an advancement for artifact generation in Microsoft 366, producing outputs that are highly cohesive, accurate, and ready for use."

This timing and posture are hard to separate from the "OpenAI-Microsoft split" rumors that dominated the prior week. Reports had surfaced that OpenAI was exploring building independent compute infrastructure beyond Azure, suggesting cracks in the partnership. But GPT-5.6's launch and full-stack integration sent a clear message at the product level: the partnership is deepening.

Yet OpenAI is simultaneously sending signals of independence. The GPT-5.6 release prominently featured endorsements from non-Microsoft enterprise customers — Stripe, Ramp, Shopify, Cisco, Balyasny Asset Management — demonstrating its standalone enterprise traction beyond Microsoft's channel. Ramp's Ian Tracey described Sol as an "end-to-end technical operator," and Shopify Senior Applied AI/ML Engineer Shane Moran noted Sol's superior intent-following in multi-stage Codex workflows.

On pricing, OpenAI continued its "efficient by default" strategy. Specific API pricing wasn't published, but the company made clear Sol delivers massive value-per-token gains — on the Artificial Analysis Intelligence Index, Sol max reasoning came within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost. Terra and Luna's positioning is unambiguous: let mid-tier and lightweight tasks achieve near-frontier performance at rock-bottom cost.

Executive Turmoil and Product Cadence: OpenAI's "Business as Usual"

The same day GPT-5.6 launched, two personnel news items cast shadows over the product celebration.

First, Fidji Simo, OpenAI's head of applied business and the company's number-two executive, announced she's stepping down from her full-time role to become a part-time advisor due to a worsening chronic neuroimmune condition. Simo wrote in an internal memo that she began medical leave three months ago and decided after recovery to prioritize her health. Second, OpenAI chief futurist Joshua Achiam announced he's leaving at the end of the month after nearly nine years — a veteran who witnessed the organization's transformation from nonprofit safety research lab to commercial giant.

The two departures have different causes — Simo's is health-related, while Achiam's exit is seen by some observers as another case of the "safety faction" bleeding out — but combined, following Ilya Sutskever's departure to found SSI and Jan Leike's move to Anthropic, the talent drain from OpenAI's safety wing has become impossible to ignore.

Yet OpenAI's choice to ship a major product like GPT-5.6 on the same day is itself a response: no matter the executive changes, the product iteration cadence continues. This "show, don't tell" approach speaks louder than any PR statement. At a critical juncture — IPO deliberations, chasing Anthropic in enterprise, exploring compute independence — OpenAI needs sustained product output to prove to investors, customers, and employees that the machine is still running at full speed.

"OpenAI's strategy is clear: when the world is debating your org chart and partnerships, shipping a model that forces every competitor into a three-month catch-up cycle is the best possible response."— Dawn Vision Editorial Assessment

Verdict: The Agentic AI Efficiency Inflection Point Is Here

Placing GPT-5.6's launch in the broader context of AI industry evolution, a clear inflection point is emerging.

2023-2024 was the "capability race" phase — competing on whose model was smarter, whose benchmark scores were higher. 2025 was the "price war" phase, as DeepSeek drove inference costs to the floor and everyone followed. Now in late 2026, the race has entered the "efficiency war" phase — not who's stronger, but who can complete the same complex tasks with fewer tokens, lower cost, and shorter time.

GPT-5.6's three-tier lineup (Sol/Terra/Luna) is a product of the efficiency war. The underlying business logic is straightforward: agentic workloads consume 10-100x more tokens than traditional chat. If per-token efficiency doesn't improve dramatically, large-scale agent commercialization is a castle in the sky. Ultra mode's multi-agent parallelism and Programmatic Tool Calling both solve the same core problem: how to make AI agents run reliably, durably, and affordably in real production environments.

For developers, GPT-5.6 sends clear signals: first, don't call the strongest model for every task — Terra and Luna already deliver good enough performance for most scenarios at massive cost savings; second, multi-agent architecture is no longer experimental — ultra mode means OpenAI has officially adopted multi-agent collaboration as the default frontier reasoning paradigm; third, design and frontend capability is the new competitive high ground — AI's leap from "writing code" to "building products" is underway.

For enterprise users, GPT-5.6's full integration into M365 Copilot means the barrier to usable AI productivity tools has dropped again. When AI can generate slides, reports, and documents that conform to corporate template standards out of the box — rather than drafts requiring heavy human revision — AI penetration in enterprises will see truly non-linear growth.

A score of 53.6 isn't the finish line. In AI's exponential iteration race, GPT-5.6's lead window may be just two to three months. But the fact that OpenAI delivered on schedule amid executive turmoil and partnership uncertainty says it all: this company's product engine is still running at full throttle.

See you tomorrow.

"GPT-5.6给人的感觉不像是一个聊天助手,更像是一个端到端的技术操作员——检查线上系统、调试问题、修改代码、验证结果、发布产物,在长会话中保持极强的上下文扎根能力。"

—— Ian Tracey,Ramp软件工程师

"GPT-5.6 felt less like a chat assistant and more like an end-to-end technical operator. It could inspect live systems, debug issues, make code changes, validate results, publish artifacts, and carry context across long sessions with strong grounding."

— Ian Tracey, Software Engineer, Ramp
GPT-5.6 · Sol · Terra · Luna · ultra模式 · 多Agent · Programmatic Tool Calling · Agents Last Exam 53.6 · Coding Agent Index 80 · M365 Copilot首选模型 · Fidji Simo离职 · Joshua Achiam离职 · 效率战争 · 设计判断力跃升
GPT-5.6 · Sol · Terra · Luna · ultra mode · multi-agent · Programmatic Tool Calling · Agents Last Exam 53.6 · Coding Agent Index 80 · M365 Copilot preferred model · Fidji Simo steps down · Joshua Achiam leaving · efficiency war · design judgment leap
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 32 个源信号生成,经编辑部人工审核。素材来源:OpenAI官方博客、TechCrunch、Satya Nadella X平台、汇通财经、36氪。

This article was generated from 32 source signals processed by the Dawn Vision cognitive engine, with editorial review. Sources: OpenAI official blog, TechCrunch, Satya Nadella on X, FX678, 36Kr.