解出答案和真正理解,是两件完全不同的事。25位全球最顶尖的数学家,刚刚联合向AI行业发出了这个警告。
9月11日,数学家陶哲轩(Terence Tao)在个人博客发布了一份声明——"A Severe Misalignment of AI in Mathematics",同步上线于mathandai.org。这份声明在一周内获得了25位菲尔兹奖得主的联署。菲尔兹奖是数学界的诺贝尔奖,每四年颁发一次,每次最多四人。25位得主联署,意味着全球近三分之一的在世菲尔兹奖得主站在了同一边。
核心指控:benchmark竞赛≠数学进步
声明的核心论点可以用一句话概括:AI实验室将数学问题作为benchmark竞赛,与数学学科的真实目标"严重错位"(severely misaligned)。
具体来说,AI实验室的做法是:选取已知有解的数学难题(如奥林匹克竞赛题、经典定理证明),训练AI模型去"攻克"它们,然后在排行榜上宣称得分提升。但数学研究的真正目标不是解出已知答案——而是提出新问题、发现新结构、建立新联系。
声明指出,这种benchmark驱动的AI研究正在产生一种危险的错觉:AI"解决"了数学问题,但并没有推动数学学科的真正进步。它更像是一个超强的考试机器,而不是一个有创造力的研究伙伴。
时间点的巧合:纳维-斯托克斯千禧年难题
这份声明的发布时间颇为耐人寻味。
就在9月8日,OpenAI刚宣布用1万个AI Agent"攻克"了纳维-斯托克斯方程千禧年难题。OpenAI称这些Agent在数小时内完成了对这一百年未解难题的"实质性推进"。这一声明在媒体上引发了巨大轰动——"AI解决了千禧年难题"的标题铺天盖地。
三天后,25位菲尔兹奖得主的声明来了。虽然没有直接点名OpenAI,但时间点的巧合让这份声明不言自明。用1万个Agent暴力搜索一个数学难题的"解",和真正理解这个难题的本质,是完全不同的两件事。
超越数学:知识生产的预警信号
声明虽然以数学为切入点,但明确指出:这个问题不限于数学。
同样的错位正在发生于所有被AI"攻克"的领域——蛋白质折叠、药物发现、材料科学、代码生成。AI实验室用benchmark分数证明AI在这些领域"超越人类",但benchmark衡量的是"解决已知问题的能力",而非"推动学科前进的能力"。
声明认为,这对更广泛的智识工作构成威胁。如果AI研究的方向被benchmark绑架,那么真正有创造性的、非标准化的知识生产将被边缘化。数学家们担心的不只是数学的未来——他们担心的是整个知识生产体系被一种"考试逻辑"所扭曲。
"解出一个已知答案是一回事,提出一个值得被解的新问题是另一回事。AI目前只擅长前者,而科学的进步依赖后者。"—— 声明核心论点
AI行业的回应压力
25位菲尔兹奖得主的联署,给AI行业带来了不小的舆论压力。
这可能是有史以来学术界对AI benchmark文化最权威、最集中的批评。菲尔兹奖得主不是反对AI——很多人在自己的研究中使用AI工具。他们反对的是将AI在数学中的价值等同于"解题得分"的单一叙事。
对AI实验室来说,这份声明提出了一个尖锐的问题:当你宣布"AI攻克了X难题"时,你是在推动科学进步,还是在制造PR素材?当你用benchmark分数来证明AI的价值时,你是在衡量真正重要的东西,还是在自说自话?
这些问题没有简单的答案。但当全球最顶尖的数学集体说"你们搞错了方向"时,AI行业至少应该停下来听一听。
明天见。
Solving an answer and truly understanding are two completely different things. Twenty-five of the world's most elite mathematicians just collectively issued that warning to the AI industry.
On September 11, mathematician Terence Tao published a statement on his blog — "A Severe Misalignment of AI in Mathematics" — simultaneously posted to mathandai.org. Within one week, the statement gathered 25 Fields Medal co-signatories. The Fields Medal is the Nobel Prize of mathematics, awarded every four years to at most four people. Twenty-five co-signatories means nearly one-third of all living Fields Medalists stood on the same side.
Core Charge: Benchmark Races ≠ Mathematical Progress
The statement's core argument in one sentence: AI labs treating math problems as benchmark competitions is "severely misaligned" with the actual goals of mathematics as a discipline.
Specifically, here's what AI labs do: select mathematically hard problems with known solutions (competition problems, classical theorem proofs), train AI models to "crack" them, then claim score improvements on leaderboards. But the real goal of mathematical research isn't solving problems with known answers — it's posing new questions, discovering new structures, establishing new connections.
The statement argues that benchmark-driven AI research is creating a dangerous illusion: AI "solves" math problems without advancing mathematics as a discipline. It's more like an extraordinarily powerful exam machine than a creative research partner.
Timing: Navier-Stokes Millennium Problem
The statement's release timing is rather intriguing.
Just three days earlier, on September 8, OpenAI announced it had used 10,000 AI agents to "crack" the Navier-Stokes equations Millennium Problem. OpenAI claimed these agents achieved "substantial progress" on this century-old unsolved problem in hours. The announcement generated enormous media buzz — headlines about "AI solving a Millennium Problem" were everywhere.
Three days later, the Fields Medal laureates' statement arrived. While it doesn't name OpenAI directly, the timing speaks for itself. Using 10,000 agents to brute-force search for a "solution" to a math problem is fundamentally different from truly understanding that problem's essence.
Beyond Math: A Warning Signal for Knowledge Production
Though framed around mathematics, the statement explicitly notes: this problem isn't limited to math.
The same misalignment is happening in every field AI has "cracked" — protein folding, drug discovery, materials science, code generation. AI labs use benchmark scores to prove AI "surpasses humans" in these areas, but benchmarks measure "ability to solve known problems," not "ability to advance a discipline."
The statement argues this threatens broader intellectual work. If AI research direction is held hostage by benchmarks, then truly creative, non-standardized knowledge production gets marginalized. The mathematicians aren't just worried about math's future — they're worried about the entire knowledge production system being warped by an "exam logic."
"Solving a known answer is one thing; posing a new question worth solving is another. AI currently excels only at the former, while scientific progress depends on the latter."— Core thesis of the statement
Pressure on the AI Industry
The co-signature of 25 Fields Medal laureates places significant public pressure on the AI industry.
This may be the most authoritative and concentrated critique of AI benchmark culture from academia in history. The Fields Medalists aren't anti-AI — many use AI tools in their own research. What they oppose is the singular narrative equating AI's value in math with "problem-solving scores."
For AI labs, the statement poses a sharp question: when you announce "AI cracked Problem X," are you advancing science or generating PR material? When you use benchmark scores to prove AI's value, are you measuring something that matters, or talking to yourselves?
These questions have no easy answers. But when the world's most elite mathematicians collectively say "you're heading in the wrong direction," the AI industry should at least stop and listen.
See you tomorrow.
解出一个已知答案是一回事,提出一个值得被解的新问题是另一回事。AI目前只擅长前者,而科学的进步依赖后者。
—— 声明核心论点
Solving a known answer is one thing; posing a new question worth solving is another. AI currently excels only at the former, while scientific progress depends on the latter.
— Core thesis of the statement
Dawn Vision, Terence Tao, Fields Medal, AI mathematics, benchmark, misalignment, knowledge production