悬念解开了。8月26日,Z.ai(智谱)正式确认,过去一周横扫多个评测榜的神秘模型 Ox Alpha——中文社区昵称"牛来"——正是 GLM 系列的新一代多模态旗舰。没有预热海报,没有发布会,用一份第三方成绩单完成了官宣。
一张成绩单的三层校准
先把数字摆出来——连同它们的来路。首发刷屏的约 80%,是开发者 Ben Davis 拿一张只有 10 道题的小样卷抽出来的读数,当时一度压过 Claude Fable 5 的约 65%、GLM-5.3 与 Grok 4.6 的约 62%;随后 113 题全量复测落在约 63%,官方成绩单定格 63.4%——对小样本的梯次领先就此作古,第一梯队的位置却站稳了。值得注意的是,这份榜单并非 Z.ai 自己选的考场——就像 OpenAI 把 Jalapeño 送去 SemiAnalysis 的公开基准一样,把成绩放在别人家的考卷上,正在成为头部实验室流行的表态方式。
商业面:两倍使用量与当天兑现的开源
Z.ai 同步给出两个商业信号:其一是模型使用量已达 DeepSeek 两倍以上——注意,这是厂商自述口径,暂无第三方审计背书;其二,开源不再是口头支票——官宣当晚权重就已推上 Hugging Face,挂在 MIT 协议下。配合消息公布当日智谱港股盘中一度上涨超 8% 的市场反应,这更像一场精心编排的组合拳:先用匿名登顶制造话题,再用认领收割注意力,最后用当场兑现的开源,一次性买断开发者生态的耐心。
为什么不办发布会
这可能是本条简报最有意思的部分。当模型能力的商品化速度快过宣传语的保鲜期,"先跑分、后认领"就成了新的发布范式:匿 名状态让评测免于人设光环的干扰,认领瞬间又自带戏剧张力。对 Z.ai 来说,这一轮组合操作的性价比,远高于一场传统 keynote。
当然也要泼半盆冷水:单一榜单领先不等于全面代差。多模态的真实生产负载远比 DeepSWE 的任务集复杂,Ox Alpha 的工程稳定性、上下文长度与成本曲线目前都是未知数。权重真正放出后的社区复现潮随即到场——113题复测与官方成绩单一起把这块金字招牌送进了压力舱,首轮结果是读数回落、闭环体面。
The suspense is over. On August 26, Z.ai (Zhipu) officially confirmed that Ox Alpha — the mysterious model that has been topping evaluation boards all week under its Chinese-community nickname "Niu Lai" ("Bull Arrives") — is the newest multimodal flagship of its GLM family. No teaser posters, no launch event; the reveal came courtesy of a third-party scorecard.
A Scoreboard and Its Three-Layer Calibration
The numbers first — along with where they came from. The viral ~80% was drawn by developer Ben Davis from a ten-question spot-check; for a moment it outshone Claude Fable 5 at about 65% and GLM-5.3 plus Grok 4.6 at about 62%. Then a full 113-question retest landed near 63%, and the official chart settled at 63.4% — the small-sample lead is history now, while the first-tier standing holds. Note whose exam this was: not one Z.ai picked itself. Just as OpenAI sent Jalapeño into SemiAnalysis's public benchmark, placing results on somebody else's answer sheet is becoming the fashionable way for frontier labs to make a statement.
The Business Side: 2x Usage and Same-Night Open Weights
Z.ai also delivered two commercial signals. First, it claims model usage has surpassed twice that of DeepSeek — vendor self-reported, with no third-party audit yet. Second, open-sourcing was no longer a promise but a fact — the weights went up on Hugging Face the same night, under an MIT license. Add the market's reaction — Zhipu's Hong Kong-listed shares briefly jumped more than 8% intraday on announcement day — and this looks like a well-choreographed combo: manufacture buzz through anonymous dominance, harvest attention by claiming authorship, then buy developer-ecosystem patience outright by shipping.
Why No Launch Event?
This may be the most interesting part of the story. When capability commoditizes faster than marketing copy expires, "benchmark first, confirm later" becomes the new release playbook: anonymity keeps evaluations free of halo effects, while the moment of claiming credit arrives with built-in drama. For Z.ai, the return on this sequence beats any traditional keynote.
A bucket of cold water is still warranted: leading one benchmark is not the same as holding a generation-wide lead. Real multimodal production workloads are far more tangled than DeepSWE's task set, and Ox Alpha's engineering stability, context length, and cost curve remain unknowns. The reproduction wave showed up the moment the weights shipped — a 113-question retest and the official chart together sent this golden brand into the pressure chamber; round one brought the reading down and closed the loop with dignity.
当评测话语权开始向公开基准转移,发新模型的正确姿势就从发布会变成了第三方成绩单。
—— Dawn Vision编辑部
When the power to evaluate shifts toward public benchmarks, the right way to launch a model turns from a keynote into a third-party scorecard.
— The Dawn Vision Editorial Desk
Ox Alpha · GLM新一代多模态 · DeepSWE三层口径: 小样本80%(10题)→复测≈63%→官方63.4% · 先跑分后认领 · 使用量DeepSeek两倍以上(厂商自述) · 权重已开源(HF/MIT) · 港股盘中+8%
Ox Alpha · New-gen GLM multimodal · DeepSWE in three layers: ~80% (10Q) -> ~63% (113Q retest) -> 63.4% official · Benchmark-first reveal strategy · Usage 2x DeepSeek (vendor-reported) · Weights live on HF under MIT · HK shares +8% intraday
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎对当日信号的交叉验证生成,核心事实经 TechCrunch 独家报道确认;DeepSWE 通过率为第三方评测数据;"使用量为 DeepSeek 两倍以上"系 Z.ai 自述口径,已在正文中明确标注。 口径更新:首发期第三方读数约80%经证实为Ben Davis 10题小样本;后续113题全量复测约63%、官方公告63.4%,三层口径已在正文分层标注。
This brief was produced by the Dawn Vision cognition engine and cross-checked against same-day signals; core facts are confirmed via TechCrunch's exclusive report. DeepSWE figures are third-party benchmark data; "over 2x DeepSeek usage" is Z.ai self-reported and flagged as such in the text. Update: the early third-party reading of ~80% proved to be Ben Davis's ten-question sample; a 113-question community retest (~63%) and the official chart (63.4%) followed, all three bases labeled separately in the text.