Focus · 焦点

GLM-5.3-Flash 当场开源
80%神话落成63.4%,与社区复测精确咬合

GLM-5.3-Flash Open-Sourced Overnight
An 80% Myth Settles Into an Audited 63.4%

漂了一周的"牛来"由智谱亲手揭盖:型号正是GLM-5.3-Flash,320B-A18B MoE、百万上下文、MIT协议,权重当晚进入Hugging Face,价格约为前代十分之一。刷屏的DeepSWE 80%呢?10道题造的神,被113题全量复测和官方成绩单一同修成了63.4%。

Zhipu pulled the label off in person: the mystery model is GLM-5.3-Flash — 320B-A18B MoE, million-token context, MIT license, weights on Hugging Face that same night, priced near a tenth of its predecessor. And the viral DeepSWE 80%? A god built on 10 questions, revised down to 63.4% by a 113-question full retest and the official chart alike.

No.044 2026.08.27 约 4 分钟阅读 ~2 min read

揭盖时刻

悬念解开了,而且解得比所有人预期的更彻底。8月26日晚,Z.ai 正式确认:过去一周横扫多个评测榜的神秘模型 Ox Alpha——中文社区昵称"牛来"——正是 GLM 系列的新一代旗舰 GLM-5.3-Flash。没有预热海报,没有发布会,甚至没有惯常那种"计划开源"的口头承诺:官宣与行动是同一个动作,权重当晚就推上了 Hugging Face,MIT 协议

规格速览

先把硬参数摆出来:3200 亿总参数、单次激活仅 180 亿(320B-A18B MoE);原生多模态架构,文本、图像、视频在同一套底座里处理,而不是外挂视觉塔;100 万 token 上下文窗口;定价约为前代的 十分之一(厂商口径)。无论下面要讲的故事有多少波折,有一个事实没有争议:Z.ai 把一个前沿规模的开放权重模型,直接放进了 Flash 价位段。

神话是怎么炼成的

复盘这条引爆链路很有必要。Ox Alpha 匿名上线一周,在多个评测位悄然登顶;随后开发者 Ben Davis 发出一篇只有 10 道题的小样本抽测,通过率约 80% 的截图点燃了社交媒体,"领先第二名约 15 个百分点"的说法随之疯转——为 agentic coding 而生的 DeepSWE 这张考卷一夜出圈。小样本不是原罪,但它天生适合造神。

全量校准

接下来发生的事,比 80% 这个数字本身好看得多:开发者随即发起 113 题全量复测,成绩落在约 63%;而智谱在官宣当晚公布的成绩单是 63.4%——与社区复测几乎精确咬合。两个细节同样重要:其一,DeepSWE v1.1 官方榜单至今没有这个模型的条目,所有数字都来自公告口径与第三方转测;其二,有审计者发现这 113 道题里有 4 题连官方参考答案本身都无法通过验证器——基准自身的成色也要打个折扣看。

排位重估

63.4% 计,GLM-5.3-Flash 在社区复测口径下坐不上 DeepSWE 的头把交椅——Claude Fable 5 约 65% 的读数仍在它前面;但它依然稳稳站在第一梯队,与 Grok 4.6、自家的 GLM-5.3(均约 62%)拉开半个身位。真正的头条不是名次,而是这次校准本身展示的双层闭环:社区科学以天为单位纠偏,厂商在同一天认领了修正后的数字。这在开源模型的历史里并不多见。

商业信号与市场反应

Z.ai 同步给出了商业面的说法:模型使用量已达 DeepSeek 两倍以上——注意,这是厂商自述、未经第三方审计的口径。资本市场则给了更即时的反馈:港股当日盘中一度上涨超 8%。"跑分营销"换来的注意力与"当场开源"兑现的开发者信任,在这一天的 K 线里短暂重合了。

冷水环节,照旧

泼冷水的时间到了:一次基准咬合不等于全面代差胜利。多模态的真实生产负载远比 DeepSWE 的任务集复杂,工程稳定性、长上下文表现、成本曲线都还要等真实流量来检验。变化在于叙事时态——叙事一直停在那句预告:"等权重真正放出之后的社区复现潮,才是这块金字招牌的第一场压力测试";现在,权重已经在场上了,压力测试从将来时变成了进行时。这也是把"牛来"请下神坛之后,这个故事最值得记住的部分。

The Unboxing Moment

The suspense ended — and resolved more decisively than anyone predicted. On the evening of August 26, Z.ai officially confirmed it: the mystery model Ox Alpha — nicknamed "Niulai" in Chinese developer circles — is GLM-5.3-Flash, the new flagship of the GLM family. No teaser poster, no keynote, none of the usual "open weights planned" hedging: announcement and action were the same gesture. The weights went up on Hugging Face that very night, under an MIT license.

The Spec Sheet, Briefly

The hard parameters first: 320 billion total parameters with just 18 billion activated per token (a 320B-A18B MoE); a natively multimodal architecture handling text, image and video on one shared backbone rather than a bolt-on vision tower; a million-token context window; and pricing set at roughly a tenth of its predecessor (vendor-reported). Whatever drama unfolds below, one fact is undisputed: a frontier-scale open-weights model has been dropped straight into the Flash price tier.

How the 80% Myth Got Built

Replaying the ignition chain matters. Ox Alpha spent a week quietly topping evaluation boards after its anonymous release; then developer Ben Davis posted a ten-question spot-check showing a pass rate of roughly 80%, and the screenshot lit social media on fire. "Fifteen points ahead of second place" went viral overnight, dragging DeepSWE — an agentic-coding benchmark few outside eval circles knew — into sudden fame. Small samples aren't a sin; they're just nature's fastest myth-factory.

Full-Sample Calibration

What happened next looks better than the 80% itself. Developers launched a full retest of all 113 questions, landing near 63%; the official chart Zhipu published on announcement night reads 63.4% — nearly perfect overlap with the community retest. Two footnotes deserve equal billing: the DeepSWE v1.1 official leaderboard never carried an entry for this model, so every number comes from announcements and third-party runs; and auditors found that 4 of the 113 reference answers cannot even pass their own verifier. The benchmark itself deserves a discount too.

Standing, Recalibrated

At 63.4%, GLM-5.3-Flash does not take the DeepSWE crown on community-retest readings — Claude Fable 5 sits ahead at roughly 65% — yet it stands firmly in tier one, half a stride clear of Grok 4.6 and its own sibling GLM-5.3 (both around 62%). The real headline is not the ranking but the calibration itself: a double closure where community science corrected course within days, and the vendor co-signed the corrected number the same night. That loop rarely closes this cleanly in open-model history.

Business Signals and the Tape

Z.ai paired the reveal with commercial claims: model usage has reportedly surpassed twice that of DeepSeek — note, vendor self-reported and third-party-unaudited. The market gave a faster verdict: Hong Kong-listed shares jumped more than 8% intraday on reveal day. For one trading session, attention bought by benchmark theater and developer trust earned by shipping weights overlapped on the same candlestick.

Cold Water, As Usual

Time for the ritual dousing: one aligned benchmark does not make a generational leap. Real multimodal workloads are messier than DeepSWE's task set; engineering stability, long-context behavior and cost curves all still owe answers to real traffic. What changed is the tense of the story — the story sat on one forward-looking promise: "the reproduction wave, once the weights ship, will be this golden brand's first stress test." The weights are on the field now, and the stress test has moved from future tense to present continuous. That, after knocking "Niulai" off its pedestal, is the part of this story worth remembering.

先跑分后认领的策略赢了一周注意力;而真正赢得体面的,是社区用113道题完成的复测——神话落地为63.4%的那一刻,纠正者与被纠正者都成了赢家。

—— Dawn Vision编辑部

Benchmark-first won a week of attention; what won respect was the community's 113-question retest — the moment the myth settled into 63.4%, correctors and corrected both came out ahead.

— The Dawn Vision Editorial Desk
GLM-5.3-Flash · Ox Alpha官方认领 · 8/26晚官宣+当场开源(HF/MIT) · 320B-A18B MoE · 百万上下文 · 价格≈前代1/10(厂商口径) · DeepSWE三层咬合: 小样本80%(Ben Davis)→社区113题≈63%→官方63.4% · v1.1榜无条目 · 审计发现4题参考答案过不了验证器 · 使用量2x DeepSeek(厂商自述) · 港股盘中+8%
GLM-5.3-Flash · Official unboxing · Announced + open-sourced same night (HF/MIT) · 320B-A18B MoE · 1M context · Price ~1/10 predecessor (vendor-reported) · DeepSWE triple alignment: 80% small sample (Ben Davis) -> ~63% community retest -> 63.4% official · No v1.1 leaderboard entry · Audit: 4 reference answers fail own verifier · Usage 2x DeepSeek (vendor-reported) · HK shares +8% intraday
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎对当日信号的交叉验证生成,核心事实经 TechCrunch 独家报道与智谱官方公告双源确认;80%/约63%/63.4% 三组数字的口径差异(10题小样本抽测 / 113题社区全量复测 / 官方公告成绩单)已在正文分层标注;使用量对比与价格比为厂商自述口径,正文中均已标记。

Produced by the Dawn Vision cognition engine and cross-checked against same-day signals; core facts are dual-sourced from TechCrunch's exclusive and Zhipu's official announcement. The differing bases of 80% (10-question sample), ~63% (113-question community retest) and 63.4% (official chart) are labeled separately in the text; usage comparison and price ratio are vendor-reported and flagged as such.