AI 商业化 · 基础设施

趋境科技半年融10亿
Token工厂模式跑通

Approaching.AI Raises 1B in 6 Months
Token Factory Model Validated

半年累计融资超10亿,日均万亿级Token产能,部分成熟业务已跨过成本线——当AI基础设施从MaaS进化到TaaS,竞争的核心从"有多少卡"变成"能产出多少高质量Token"。

Over 1 billion yuan raised in six months, trillions of tokens produced daily, some mature businesses already profitable — when AI infrastructure evolves from MaaS to TaaS, the core of competition shifts from "how many GPUs you have" to "how many high-quality tokens you can produce."

No.013 2026.07.14 约 5 分钟阅读 ~5 min read

半年,10亿。

7月13日,趋境科技(Approaching.AI)正式宣布完成A轮融资。半年内累计融资超过10亿元,本轮由河南投资集团汇融基金领投,老股东全部超额跟投。

10亿人民币的融资规模,在当前的AI基础设施赛道不算最大的——CoreWeave上市时市值几百亿美元,Nebius也融了几十亿。但趋境科技的故事有意思的地方在于:它不是又一家"卖GPU算力"的云服务商,它提出了一个新概念——Token as a Service(TaaS,Token即服务)

简单说:以前你买的是"GPU小时",现在你买的是"高质量Token"。听起来好像只是换了个说法,但背后的商业逻辑完全不同。

从MaaS到TaaS:卖的不是算力,是结果

过去两年,AI基础设施的主流模式是MaaS(Model as a Service)——我搭好平台、接入一堆模型,你按Token量调用付费。OpenAI、Anthropic、各家云厂商,都是这个模式。

但MaaS有个问题:客户真正需要的不是"调用了多少Token",而是"完成了多少业务结果"。你花了100万调用API,结果生成的内容质量参差不齐、经常出错、速度时快时慢——这100万花得值不值?没人说得清。

趋境科技的TaaS模式,就是把这个问题翻过来:我不卖给你算力,也不卖给你模型,我卖给你"高品质AI Token"。什么叫高品质?就是在高并发下依然能保持:低首Token响应时延、稳定高输出速度、持续输出质量、可靠的结构化输出和函数调用、可控的单位成本。

这几个指标单独拿出来都不难,但要在真实生产负载下同时满足,难度就指数级上升了。不同的能力组合,生产效率可能差几倍甚至几十倍。趋境科技的核心竞争力,就是把这些指标在真实生产环境里同时跑通,而且能规模化复制。

数据是最好的证明:从2026年春节以来,趋境科技单台算力的Token生产效率提升了3倍以上,高品质Token总产能增长超过30倍。某头部万亿级参数大模型,高品质日均产量已经稳定突破万亿量级。更重要的是:部分成熟业务已经跨过成本线

"跨过成本线"这五个字,在AI基础设施行业里的分量,怎么说都不为过。现在做推理服务的公司一大堆,但真正盈利的没几家——大家都在烧钱抢市场、拼规模、比谁的价格低。趋境科技说自己部分业务已经盈利,这就很说明问题了。

"少模型、深优化":反共识的技术路线

趋境科技的技术路线也很有意思,叫"少模型、深优化"——和行业主流的"多模型、大而全"路线正好相反。

现在的MaaS平台,比拼的是接入了多少模型——OpenAI的、Anthropic的、Google的、开源的,越多越好,最好是一键切换。但趋境科技不这么干。他们就盯着少数几个真正有生产需求的大模型,往死里优化:模型切分、显存管理、异构协同、缓存复用、故障恢复、弹性扩缩容——能优化的地方全都优化一遍。

为什么?因为企业级客户最终为业务结果付费,不是为模型兼容数量付费。你接入100个模型,但每个模型都跑得又慢又贵又不稳定,有什么用?不如把两三个真正用得上的模型优化到极致。

这个逻辑听起来很朴素,但在行业里反而是少数派。为什么?因为"多模型"好做PPT、好讲故事、好融资,"深优化"又苦又累又不出活,短期看不到成效。但真正到了生产环境,客户用脚投票的时候,"深优化"的价值就体现出来了。

趋境科技的融资故事,本质上是在讲一件事:AI基础设施的竞争,已经从"谁的盘子大"进入"谁的效率高"的阶段了。前两年大家都在抢地盘、铺算力、堆规模,现在潮水退了,才知道谁在裸泳。能把单位算力的Token产出做到最高、能把成本线打穿、能让客户真正用起来的公司,才是能活到最后的。

Token工厂的时代,才刚刚开始。

明天见。

Six months. One billion yuan.

On July 13, Approaching.AI (趋境科技) officially announced its Series A round. Cumulative funding over the past six months exceeds 1 billion yuan, led by Henan Investment Group's Huirong Fund, with all existing shareholders over-subscribing.

A 1 billion yuan raise isn't the biggest in the current AI infrastructure race — CoreWeave went public at a market cap of tens of billions; Nebius has raised billions too. But what makes Approaching.AI's story interesting is that it's not just another cloud provider "selling GPU hours." It's proposing a new concept: Token as a Service (TaaS).

Put simply: before, you bought "GPU hours." Now, you buy "high-quality tokens." It sounds like just rebranding, but the underlying business logic is entirely different.

From MaaS to TaaS: Selling Results, Not Compute

For the past two years, the dominant AI infrastructure model has been MaaS (Model as a Service) — I set up the platform, integrate a bunch of models, and you pay by token volume. OpenAI, Anthropic, every cloud provider — all follow this model.

But MaaS has a problem: what customers actually need isn't "how many tokens they've called" — it's "how many business outcomes they've achieved." You spend a million on API calls, but the generated content quality is all over the place, errors are frequent, speed fluctuates — was that million well spent? Nobody really knows.

Approaching.AI's TaaS model flips this: I don't sell you compute, I don't sell you models — I sell you "high-quality AI tokens." What does "high-quality" mean? Under high concurrency, it still maintains: low first-token latency, stable high output speed, consistent output quality, reliable structured output and function calling, and controllable unit cost.

None of these metrics are particularly hard on their own. But achieving all of them simultaneously under real production loads? The difficulty rises exponentially. Different combinations of capabilities can produce efficiency gaps of several times — even tens of times. Approaching.AI's core competitive edge is making all these metrics work together in real production environments, and doing it at scale.

The numbers speak for themselves: since the 2026 Spring Festival, Approaching.AI has increased per-GPU token production efficiency by more than 3x, and total high-quality token capacity has grown over 30x. For one leading trillion-parameter model, daily high-quality output has steadily broken through the trillion level. Most importantly: some mature businesses have already crossed the profitability line.

"Crossed the profitability line" — those five words carry more weight in the AI infrastructure industry than almost anything else. There are tons of companies doing inference services today, but very few are actually profitable — everyone's burning cash to grab market share, racing for scale, undercutting each other on price. When Approaching.AI says parts of its business are already profitable — that says a lot.

"Few Models, Deep Optimization": A Counter-Consensus Technical Route

Approaching.AI's technical approach is also interesting. They call it "few models, deep optimization" — exactly the opposite of the industry mainstream "many models, one-stop-shop" approach.

Today's MaaS platforms compete on how many models they've integrated — OpenAI, Anthropic, Google, open-source ones, the more the better, ideally with one-click switching. But Approaching.AI doesn't play that game. They focus on just a handful of models with real production demand, and optimize them to death: model sharding, VRAM management, heterogeneous coordination, cache reuse, failover, elastic scaling — everything that can be optimized, they optimize.

Why? Because enterprise customers ultimately pay for business outcomes, not for the number of compatible models. You can integrate 100 models, but if each one is slow, expensive, and unreliable — what's the point? Better to optimize two or three models people actually use to the absolute maximum.

This logic sounds straightforward, but it's actually a minority position in the industry. Why? Because "many models" makes for good slides, good stories, and good fundraising. "Deep optimization" is hard, tedious, and doesn't show quick results — short-term payoff is invisible. But when you get to production environments and customers vote with their feet, the value of "deep optimization" becomes apparent.

Approaching.AI's funding story is essentially about one thing: AI infrastructure competition has moved from "who has the biggest plate" to "who has the highest efficiency" phase. The past two years everyone was grabbing territory, deploying compute, piling up scale. Now the tide is going out, and we'll see who's been swimming naked. The companies that maximize token output per unit of compute, that break through the cost line, that customers actually use — those are the ones that survive.

The era of the Token Factory is just beginning.

See you tomorrow.

企业级客户最终为业务结果付费,不是为模型兼容数量付费。

—— 趋境科技团队

Enterprise customers ultimately pay for business outcomes, not for the number of compatible models.

— Approaching.AI Team
趋境科技 · Approaching.AI · Token工厂 · TaaS · AI基建 · 推理优化 · 10亿融资 · 少模型深优化
Approaching.AI · token factory · TaaS · AI infrastructure · inference optimization · 1B yuan funding · few models deep optimization
Sources · 信源 Sources

本文基于 Dawn Vision 认知引擎处理的 12 个源信号生成,经编辑部人工审核。素材来源:量子位、36氪、人民财讯。

This article was generated by the Dawn Vision cognitive engine processing 12 source signals, with human editorial review. Sources: QbitAI, 36Kr, People's Finance News.