大模型

六月模型
军备赛

The June Model
Arms Race

GPT-5.6、Claude 4.8、Gemini 3.5 Pro 罕见同一时间窗口发布旗舰模型,150万Token上下文、科学推理登顶、国产模型市值破万亿——史上最密集的模型军备赛正在上演。

GPT-5.6, Claude 4.8, Gemini 3.5 Pro unusually released flagships in the same time window: 1.5M Token context, science reasoning crown, domestic models breaking trillion-yuan market cap -- the densest model arms race in history is underway.

No.002 2026.06.25 约 6 分钟阅读 ~6 min read

2026 年 6 月,AI 实验室像是约好了一样。

OpenAI、Google、Anthropic 三家顶级实验室罕见地选择了同一个时间窗口发布旗舰模型:Claude Sonnet/Opus 4.8 已上线,GPT-5.6 即将全面开放(上下文从 5.5 的 100 万跃升到 150 万 Token),Gemini 3.5 Pro 推理准确率提升 35%。上一次这么密集的发布,还要追溯到 2023 年 GPT-4 刚发布的时候——但那时候只有一个玩家,这次是三家同时出牌。

三家旗舰,各有杀招

先看成绩单:Claude Opus 4.8 以 76.4 分的 ScienceQA 成绩登顶科学推理王座。这是一个很有意思的信号——Anthropic 选择在"严谨性"这个维度建立护城河,而不是拼上下文长度或者多模态能力。对于需要写论文、做研究、分析复杂数据的用户来说,Claude 正在成为首选。

GPT-5.6 的杀手锏是上下文窗口:150 万 Token。这是什么概念?你可以把一整套《红楼梦》+《三国演义》+《水浒传》+《西游记》塞进去,还有余量。对于需要处理超长文档、分析整个代码库、阅读整本财报的场景,这是质的飞跃。

Gemini 3.5 Pro 则在多模态推理上继续领跑——毕竟谷歌搜索+YouTube+安卓的生态数据,是另外两家短期内难以复制的优势。视频理解、实时翻译、跨模态推理,Gemini 仍然是地表最强。

"当三家都在同一个月发布旗舰,说明模型能力的差距正在以周为单位缩短。" —— AI 行业分析师

开源阵营:战火同样激烈

闭源打得火热,开源也没闲着。

DeepSeek V4、Qwen3 全家桶、Llama 4、Gemma 4 轮番上阵。国产模型这边,智谱 GLM-5.2 在代码能力上超越多款海外主流模型,智谱港股市值突破万亿港元——上市半年涨了 18 倍。

更值得关注的是阿里通义千问发布的 Qwen-AgentWorld——全球首款原生语言世界模型。单一底座兼容代码终端、网页、手机、桌面 OS 等七大交互环境,多环境协同任务得分超越了 GPT-5.4 和 Claude Opus。这意味着什么?意味着国产模型不再只是"跟随者",在 Agent 原生交互这个前沿赛道上,中国团队已经跑到了前面。

还有一个出人意料的玩家:Cursor。这家以 AI IDE 闻名的公司,在 Compile 开发者大会上发布了从零自研的通用大模型——不是基于开源基座微调,是真的从零训练。由 SpaceX 提供 GPU 集群支撑训练,能力从代码生成扩展到了文档分析、项目统筹等通用任务。当一个 IDE 公司开始自研通用大模型,说明"模型即产品"的时代已经到来。

军备赛的终局是什么?

但模型越来越强,用户的感知却越来越弱。

GPT-4 发布时,全世界为之震动。GPT-5 发布时,大家觉得"嗯,确实变强了"。到了 GPT-5.6,普通用户可能根本说不出来它和 5.5 有什么区别。模型能力的边际收益正在递减——就像手机芯片从 3nm 进化到 2nm,参数党很兴奋,但普通用户刷抖音还是一样流畅。

真正的竞争正在从"模型有多强"转向"你能用模型做什么"。MCP 协议、Agent 编排、工作流集成、垂直场景落地——这些"脏活累活"正在成为新的护城河。一个用着 Claude 3.5 但精通 Prompt Engineering 和 Agent 编排的人,可能比一个用着 GPT-5.6 但只会聊天的人效率高十倍。

这也是为什么 Dawn Vision 反复强调:不要纠结哪个模型"最聪明",要关注你的工作流有没有被 AI 重构。模型是工具,工具再锋利,不会用也是白搭。

六月之后,看点是什么?

这场军备赛还远没有结束。下半年值得关注的几个节点:

一是多模态的真正融合——不是文字配图片,而是模型能像人一样在视觉、听觉、文字之间无缝切换理解。二是 Agent 能力的标准化——当所有模型都能调用工具,谁的 Agent 更可靠、更可控将成为关键。三是端侧模型的爆发——手机、PC、IoT 设备上运行的小模型,可能会重新定义 AI 的使用场景。

但对于普通用户来说,最好的消息是:竞争越激烈,价格越便宜。当三家顶级实验室和无数开源模型打得头破血流,最终受益的是每一个用 AI 的人。

模型会越来越强,价格会越来越低,门槛会越来越平。你需要做的,就是别在这场军备赛里当观众——下场用起来。


明天见。

In June 2026, AI labs acted like they had an agreement.

OpenAI, Google, and Anthropic -- three top labs -- unusually chose the same time window to release flagship models: Claude Sonnet/Opus 4.8 is live; GPT-5.6 is opening fully (context jumping from 5.5's 1 million to 1.5 million Tokens); Gemini 3.5 Pro inference accuracy improved 35%. The last time releases were this dense was when GPT-4 launched in 2023 -- but back then there was only one player; this time all three are playing their cards simultaneously.

Three Flagships, Each With a Killer Move

First, the report card: Claude Opus 4.8 claimed the science reasoning throne with a ScienceQA score of 76.4. That's an interesting signal -- Anthropic chose to build a moat in the "rigor" dimension, rather than competing on context length or multimodal capability. For users writing papers, doing research, or analyzing complex data, Claude is becoming the first choice.

GPT-5.6's killer feature is context window: 1.5 million Tokens. What does that mean? You can fit the entire "Dream of the Red Chamber" + "Romance of the Three Kingdoms" + "Water Margin" + "Journey to the West" in there and still have room. For scenarios requiring ultra-long document processing, entire codebase analysis, or reading full financial reports, it's a qualitative leap.

Gemini 3.5 Pro continues to lead in multimodal reasoning -- after all, Google Search + YouTube + Android ecosystem data is an advantage the other two can hardly replicate short-term. Video understanding, real-time translation, cross-modal reasoning; Gemini remains the strongest on the planet.

"When all three release flagships in the same month, it means gaps in model capability are narrowing on a weekly basis." -- AI industry analyst

Open-Source Camp: Equally Fierce Fighting

Closed-source is red-hot; open-source isn't idle either.

DeepSeek V4, the Qwen3 family, Llama 4, Gemma 4 are all taking turns. On the domestic model front, Zhipu GLM-5.2 surpassed multiple overseas mainstream models in code capability, and Zhipu's Hong Kong stock market cap broke a trillion Hong Kong dollars -- up 18x in the six months since listing.

More noteworthy is Alibaba Tongyi Qwen's release of Qwen-AgentWorld -- the world's first native language world model. A single base supports seven interaction environments including code terminals, web, mobile, and desktop OS; multi-environment collaborative task scores exceed GPT-5.4 and Claude Opus. What does this mean? Chinese models are no longer just "followers"; on the frontier track of Agent-native interaction, Chinese teams have pulled ahead.

There's also an unexpected player: Cursor. The company known for its AI IDE released a general-purpose LLM trained from scratch at the Compile developer conference -- not fine-tuned from an open-source base, but truly trained from zero. Backed by SpaceX GPU clusters for training, its capabilities expanded from code generation to document analysis, project coordination, and other general tasks. When an IDE company starts training its own general-purpose LLM from scratch, it means the "model-as-product" era has arrived.

What Is the Endgame of the Arms Race?

But as models get stronger, users perceive less difference.

When GPT-4 launched, the world was shaken. When GPT-5 launched, people felt "yeah, definitely stronger." By GPT-5.6, ordinary users might not be able to tell how it differs from 5.5 at all. Marginal returns on model capability are diminishing -- just like phone chips evolving from 3nm to 2nm; spec nerds get excited, but regular users scrolling TikTok experience the same smoothness.

Real competition is shifting from "how strong is your model" to "what can you do with the model." MCP protocols, Agent orchestration, workflow integration, vertical scenario landing -- these "dirty jobs" are becoming new moats. Someone using Claude 3.5 but mastering Prompt Engineering and Agent orchestration might be ten times more effective than someone using GPT-5.6 who only knows how to chat.

That's also why Dawn Vision repeatedly emphasizes: don't obsess over which model is "smartest"; focus on whether your workflow has been restructured by AI. Models are tools; no matter how sharp the tool, it's useless if you can't wield it.

What to Watch After June?

This arms race is far from over. Key nodes to watch in H2:

First, true multimodal fusion -- not text with images, but models seamlessly switching understanding across vision, audio, and text like humans. Second, Agent capability standardization -- when all models can call tools, whose Agents are more reliable and controllable becomes key. Third, on-device model explosion -- small models running on phones, PCs, and IoT devices may redefine AI use cases.

But for ordinary users, the best news is: the fiercer the competition, the cheaper the prices. When three top labs and countless open-source models fight tooth and nail, everyone using AI ultimately benefits.

Models keep getting stronger, prices keep falling, barriers keep leveling. What you need to do is stop being a spectator in this arms race -- get in and use them.


See you tomorrow.

当三家都在同一个月发布旗舰,说明模型能力的差距正在以周为单位缩短。

—— AI 行业分析师
2026 年中模型能力对比 · 开源模型生态全景 · 国产大模型突围路径
Mid-2026 model capability comparison, open-source model ecosystem panorama, Chinese LLM breakout path
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 12 个源信号自动生成,经编辑部人工审核。素材来源包括:各模型发布信息、HuggingFace 榜单、智谱市值数据、Qwen-AgentWorld 技术报告、Cursor Compile 大会。

Auto-generated by Dawn Vision's cognitive engine from 12 source signals, editorially reviewed. Sources include: model launch information, HuggingFace leaderboards, Zhipu market cap data, Qwen-AgentWorld technical report, Cursor Compile conference.