朋友们,B站搞了个大活——AI无限竞技场,9月16日正式上线。简单说就是:100多个AI模型排排坐,用户跟它们聊天,聊完投票,谁回答好选谁。最后看排名。
结果出来了:GPT-6 Astra第一名,GLM-5.3第二名,Claude Fable 5.1第三名。前五名里,国产模型占了三席。
这个排名有点意思
你品品这个画面:OpenAI最强的模型GPT-6拿了第一,这不意外——人家砸了几百亿美元训练出来的,不拿第一说不过去。
但有意思的是GLM-5.3力压Claude排到了第二。Claude什么水平?Fable 5.1是Anthropic的最新旗舰,在硅谷的口碑一直很硬。结果到了B站用户的手里,被智谱的GLM给超了。你可以说这是因为中文场景国产模型更懂中文用户,但"更懂用户"本身就是一种竞争力——技术benchmark上赢不算赢,用户投票赢才是真赢。
1.9亿月活的AI审美
更值得关注的是背后的用户数据。B站AI内容月活用户1.9亿,AI内容消费时长同比增长72%。将近2亿人每个月在B站上看AI相关内容——这个用户基盘,比很多AI产品的月活都大。
当1.9亿人用投票的方式告诉你"哪个AI更好用",这个数据的价值就远超任何技术评测。因为benchmark测的是标准化任务,而用户投票测的是真实世界里的主观满意度。一个模型在代码生成上得99分没用,如果用户觉得它"说话像机器人"就不会投它。
实用提醒
第一,别只看第一名。竞技场的价值在于不同场景下的投票分布——如果你主要用AI写代码,可能Claude更适合你;如果你用AI写中文文案,GLM可能体验更好。排行榜是平均值,不是你的最优解。
第二,B站这个竞技场的商业模式值得关注。1.9亿月活的AI偏好数据,对AI厂商来说是金矿。未来如果B站开放"模型推荐"或者"AI产品导流",那就是一个全新的AI分发渠道——比应用商店更精准,比搜索引擎更贴近使用场景。
第三,去投一票。别光看别人评,自己上手试试,投出你心中的最佳AI。毕竟,1.9亿人的集体审美,比任何专家评测都靠谱。
今天就槽到这里,明天继续。
Friends, Bilibili just dropped something big — the AI Infinite Arena, officially launched September 16. The concept: 100+ AI models sit in a row, users chat with them, vote for the better answer, and rankings emerge in real time.
Results are in: GPT-6 Astra takes first place, GLM-5.3 second, Claude Fable 5.1 third. Of the top five, domestic models hold three seats.
This Ranking Is Interesting
Let that picture sink in: OpenAI's most powerful model, GPT-6, takes the crown — no surprise there. You spend hundreds of billions training it, it better win.
But what's interesting is GLM-5.3 leapfrogging Claude into second place. Claude Fable 5.1 is Anthropic's latest flagship — the one with sterling reputation in Silicon Valley. And yet, in the hands of Bilibili users, it got overtaken by Zhipu's GLM. You could argue this is because domestic models "get" Chinese users better in Chinese-language scenarios. But "understanding users better" is itself a competitive advantage — winning on a technical benchmark doesn't count. Winning user votes does.
190 Million Monthly Users' AI Taste
The user data behind this is even more telling. Bilibili's AI content has 190 million monthly users, with AI content watch time up 72% year-over-year. Nearly 200 million people watching AI content on Bilibili every month — that user base is larger than the MAU of many standalone AI products.
When 190 million people vote on "which AI is better," that data's value dwarfs any technical evaluation. Benchmarks test standardized tasks; user votes test real-world subjective satisfaction. A model scoring 99 on code generation doesn't matter if users think it "talks like a robot" and refuse to vote for it.
Practical Reminders
First, don't just look at the winner. The Arena's value lies in vote distribution across different scenarios — if you mainly use AI for coding, Claude might suit you better; for Chinese copywriting, GLM might deliver a superior experience. The leaderboard is an average, not your personal optimum.
Second, Bilibili's arena business model is worth watching. AI preference data from 190 million monthly users is a gold mine for AI companies. If Bilibili eventually opens "model recommendations" or "AI product referrals," that's an entirely new AI distribution channel — more precise than app stores, closer to usage context than search engines.
Third, go vote. Don't just watch others evaluate — try it yourself and cast your ballot for your best AI. After all, the collective taste of 190 million people is more reliable than any expert evaluation.
That's all the roasting for today. More tomorrow.
技术benchmark上赢不算赢,用户投票赢才是真赢。1.9亿人的集体审美,比任何专家评测都靠谱。
—— Dawn Vision编辑部
Winning on a technical benchmark doesn't count. Winning user votes does. The collective taste of 190 million people is more reliable than any expert evaluation.
— The Dawn Vision Editorial Desk
别只看排行榜第一名——竞技场的价值在于不同场景下的投票分布;1.9亿月活的AI偏好数据对AI厂商是金矿,关注B站未来可能的AI分发商业模式;去自己投一票,集体审美比专家评测更靠谱。
Don't just look at the leaderboard winner — the Arena's value is in vote distribution across scenarios; 190M MAU's AI preference data is a gold mine for AI companies — watch for Bilibili's future AI distribution model; go vote yourself — collective taste beats expert evaluation.
Dawn Vision, B站, bilibili, AI竞技场, GPT-6, GLM-5.3, Claude Fable, 大模型评测, AI内容, 1.9亿月活, 用户投票
Dawn Vision, Bilibili, AI Arena, GPT-6, GLM-5.3, Claude Fable, LLM evaluation, AI content, 190M MAU, user voting
Sources · 信源 Sources
本文基于 Dawn Vision 认知引擎处理的 8 个源信号生成,经编辑部人工审核。素材来源:IT之家、B站官方。
This article was generated by the Dawn Vision cognitive engine processing 8 source signals, with human editorial review. Sources: IT之家, Bilibili Official.