大模型 · 商业分析

Gemini跌出全球前十
三连发Flash但Pro难产

Gemini Drops Out of Top 10
Three Flash Models, No Pro

7月21日谷歌连发3.6 Flash、3.5 Flash-Lite、3.5 Flash Cyber三款模型,输出token效率提升17%,但旗舰Gemini 3.5 Pro仍在“与合作伙伴测试中”,Arena.ai榜单上谷歌无一模型进入前十。

On July 21 Google released three models — 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber — with 17% better output token efficiency, but flagship Gemini 3.5 Pro remains “testing with partners,” and not a single Google model makes the top ten on Arena.ai.

No.020 2026.07.23 约 5 分钟阅读 ~5 min read

Transformer是谷歌发明的。

2017年那篇《Attention Is All You Need》论文的作者大多在Google Brain。但2026年7月的第三方评测榜单上,谷歌没有任何一款模型进入全球前十。7月21日,Google DeepMind一天内发布了Gemini 3.6 Flash、Gemini 3.5 Flash-Lite、Gemini 3.5 Flash Cyber三款模型,却唯独不见曾经承诺的旗舰Gemini 3.5 Pro——公告里的措辞是“目前正在与合作伙伴测试,准备好后将尽快广泛发布”。

Flash三连发:便宜、快、好用,但不是旗舰

先看新产品本身。客观说,这三款Flash模型确实有亮点。

Gemini 3.6 Flash是主力型号,输入价格$1.50/百万token、输出$7.50/百万token,比3.5 Flash更便宜。在Artificial Analysis Index上,3.6 Flash比3.5 Flash少用17%的输出token,在某些基准(如Datacurve的DeepSWE)上观察到高达65%的效率提升。编码能力从DeepSWE的37%提升到49%,MLE Bench从49.7%提升到63.9%,OSWorld-Verified从78.4%提升到83.0%。

3.5 Flash-Lite是性价比杀手:输入仅$0.30/百万token,输出$2.50/百万token,输出速度达350 token/s,在SWE-Bench Pro(54.2% vs 49.6%)和OSWorld-Verified(74.0% vs 65.1%)上甚至超过了Gemini 3 Flash。3.5 Flash Cyber是面向安全领域的专用模型,通过CodeMender安全Agent提供服务,仅供政府和可信合作伙伴使用。

这些数据很好看。但问题是:这都是“Flash”系列,主打效率和性价比,不是最强模型。在前沿模型竞赛中,大家比的是Pro/Ultra级别的旗舰模型——谁的推理最强、谁的多模态最准、谁的Agent能力最可靠。在这些维度上,谷歌已经很久没有拿出令人信服的答案了。

旗舰跳票的背后:谷歌AI的结构性困境

Gemini 3.5 Pro原本承诺在2026年6月发布。现在已经是7月底,官方措辞变成了“仍在与合作伙伴测试中”。这不是第一次谷歌旗舰模型跳票——Gemini Ultra在发布后也经历了漫长的延期和能力争议。

为什么一家发明了Transformer、拥有DeepMind这样顶级研究机构、年AI投入数百亿美元的公司,在旗舰模型竞赛中持续落后?几个结构性因素值得注意。

第一,谷歌的“研究文化”与“产品文化”始终存在张力。DeepMind的科学家想发论文、做基础研究;Google Cloud和产品团队需要稳定可靠的模型服务客户。两者之间的协调成本在大模型时代被急剧放大。OpenAI是一家“产品驱动研究”的公司——研究的目标就是产品化。谷歌更像是“研究驱动产品”——有了研究突破再想怎么包装成产品,这个节奏天然慢半拍。

第二,Flash优先的战略反映了谷歌对AI商业化的务实判断。在大多数企业场景中,用得最多的不是最强的模型,而是性价比最高的模型。谷歌云在企业市场需要的是能大规模部署、成本可控、延迟够低的模型,Flash系列精准满足了这个需求。但长期只靠Flash打天下,会让谷歌在“前沿能力标杆”这个心智高地上持续失分。

第三,谷歌正在同时训练Gemini 4。DeepMind在公告中明确提到”我们最雄心勃勃的预训练运行——Gemini 4已经启动“。这意味着3.5 Pro的资源可能被部分抽调到了4.0的开发上。如果Gemini 4能在年底前发布并带来代际级别的提升,现在的等待可能是值得的;但如果继续跳票,谷歌在AI竞赛中的位置会更加尴尬。

明天见。

Google invented the Transformer.

Most authors of that 2017 “Attention Is All You Need” paper were at Google Brain. But on July 2026 third-party leaderboards, not a single Google model places in the global top ten. On July 21, Google DeepMind shipped three models in one day — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3.5 Flash Cyber — yet the promised flagship Gemini 3.5 Pro was conspicuously absent, with the announcement noting only that it's “currently testing with partners and we plan to make it broadly available as soon as it's ready.”

A Triple Flash Release: Cheap, Fast, Good — But Not Flagship

Let's give the new products their due. Objectively, these three Flash models have real highlights.

Gemini 3.6 Flash is the workhorse, priced at $1.50/1M input tokens and $7.50/1M output tokens — cheaper than 3.5 Flash. On the Artificial Analysis Index, 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash; on some benchmarks like Datacurve's DeepSWE, efficiency gains of up to 65% were observed. Coding jumped from 37% to 49% on DeepSWE; MLE Bench rose from 49.7% to 63.9%; OSWorld-Verified improved from 78.4% to 83.0%.

3.5 Flash-Lite is the value killer: just $0.30/1M input and $2.50/1M output, running at 350 tokens/s, and on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%) it even outperforms Gemini 3 Flash. 3.5 Flash Cyber is a security-specialized model delivered through the CodeMender agent, available exclusively to governments and trusted partners.

The numbers look good. But here's the thing: these are all “Flash” tier, built for efficiency and cost-effectiveness, not maximum capability. In the frontier model race, the competition is over Pro/Ultra-class flagships — whose reasoning is strongest, whose multimodal is sharpest, whose agents are most reliable. On those dimensions, Google hasn't delivered a convincing answer in quite some time.

Behind the Flagship Delay: Google AI's Structural Dilemma

Gemini 3.5 Pro was originally promised for June 2026. It's now late July, and the official language is “still testing with partners.” This isn't the first time a Google flagship has slipped — Gemini Ultra also went through extended delays and capability controversies after launch.

Why does a company that invented the Transformer, houses a research organization as elite as DeepMind, and spends tens of billions annually on AI keep falling behind in the flagship race? Several structural factors stand out.

First, tension between Google's “research culture” and “product culture” persists. DeepMind scientists want to publish papers and do fundamental research; Google Cloud and product teams need stable, reliable models to serve customers. Coordination costs between these two cultures have been magnified dramatically in the LLM era. OpenAI is a “product-driven research” company — research's goal is productization. Google operates more like “research-driven products” — package research breakthroughs into products after they happen, a rhythm that naturally lags half a beat behind.

Second, the Flash-first strategy reflects a pragmatic bet on AI commercialization. In most enterprise scenarios, the most-used model isn't the smartest one — it's the one with the best cost-performance ratio. What Google Cloud needs in the enterprise market are models that can deploy at scale with predictable costs and low latency; the Flash line precisely serves this need. But relying solely on Flash over the long run concedes the “frontier capability benchmark” mindshare continuously.

Third, Google is already training Gemini 4. DeepMind explicitly noted in the announcement that “our most ambitious pre-training run yet, for Gemini 4, has started.” That means some resources originally slated for 3.5 Pro may have been redirected to 4.0. If Gemini 4 ships before year-end with generational improvements, the wait might prove worthwhile; if it slips again, Google's position in the AI race will become even more awkward.

See you tomorrow.

发明Transformer的公司,在大模型排行榜上排不进前十——这大概是2026年AI行业最讽刺的画面之一。

—— Dawn Vision编辑部

The company that invented the Transformer can't crack the top ten on LLM leaderboards — one of the more ironic images in the 2026 AI industry.

— The Dawn Vision Editorial Desk
Gemini 3.6 Flash · Flash-Lite · Flash Cyber · 3.5 Pro延期 · 大模型排名 · 谷歌AI战略 · DeepMind · Arena.ai
Gemini 3.6 Flash · Flash-Lite · Flash Cyber · 3.5 Pro delayed · LLM rankings · Google AI strategy · DeepMind · Arena.ai
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 10 个源信号生成,经编辑部人工审核。素材来源:Google DeepMind博客、搜狐科技、新浪财经。

This article was generated by the Dawn Vision cognitive engine processing 10 source signals, with human editorial review. Sources: Google DeepMind Blog, Sohu Tech, Sina Finance.