Moxing AI raises a fresh round and sells trillions of tokens a day as inference factories boom

On 10 August, Moxing AI, a “token super-factory” founded over two years ago, closed another financing. It announced an A round led by Yida Capital, with several strategic resource partners and returning shareholders following on, some over-allocated. That came only three months after its nine-figure yuan Pre-A in May.

Moxing AI token super-factory
Moxing AI’s token super-factory. Source: Zhidx

Two rounds in three months is rare in today’s funding climate. Rarer is the company’s scale. In two years its average daily token calls reached the trillions level, with customers and partners spanning internet firms, large model houses and industry leaders.

Capital is chasing a market whose barrier looks like it is falling. DeepSeek, Kimi, GLM and MiniMax have shipped open models densely. Firms can download weights directly, and generic inference frameworks mature. Some claim model inference is becoming a “anyone with cards can do it” business.

But that is not so. “Models are open, but the efficient deployment way is not all open,” Moxing AI founder and CEO Xu Lingjie tells Zhidx. A model running only means tokens are produced. Whether, under heavy concurrency, you can hold latency, throughput, success rate and cost decides if those tokens can be sold.

Moxing AI’s disclosed daily trillions of tokens are actual metered usage that already generates revenue, not theoretical capacity. On current growth, 2026 revenue is expected at the several-hundred-million yuan level. Some models are already break-even. The whole business is nearing scaled profitability. Its market tactic is to serve large B customers first.

Open models expanded the market for token factories and redrew the division of labour. The easier weights are to get, the less scarce mere “owning a model” becomes. Capital shifted attention to the production step after the model, to who can run the same model faster, steadier and cheaper.

The next target is daily output of 10 trillion tokens. Xu, whose career runs from Nvidia and AMD through Alibaba Cloud GPU infrastructure to Biren, spent twenty years at the cross of chip, system and cloud. CTO Jin Chen brings high-performance computing and inference optimisation. Their edge is full-stack system optimisation: separating prefill and decode, tuning memory and load, raising cache hit rate and adapting many chips.

The second barrier is the super-node. As models and agent tasks grow, an eight-card server hits memory and interconnect ceilings. Super-nodes link more accelerators through high-bandwidth, low-latency interconnect for larger parallel work. Jin says human interaction at 100 to 200 tokens a second already feels smooth, but complex agents may demand 1,000 a second. Moxing AI already runs some models at hundreds per second. The push to 1,000 is under development. Bigger interconnect brings bigger “blast radius” when one chip fails, so fault isolation and tolerance matter as much as peak performance.

Open-source model prosperity lowered the entry to AI apps and pushed competition into deeper infrastructure. Moxing AI has passed first-stage validation. The real test comes as daily volume heads to 10 trillion and model, hardware and load types multiply. Whether it can keep success rate, latency, cost and stability will decide if the commercial loop keeps scaling.

Editor’s note: This is an adapted translation of the original Zhidx report. It has been trimmed and restructured for readability for an international business audience.

Leave a comment