Moxing AI Token factory raises again, sells trillions daily

Moxing AI, the “Token super factory” founded just over two years ago, has closed another funding round. The company announced a Series A led by Yida Capital, with several strategic resource partners joining, existing shareholders adding more, and some over-allocating. The round follows a several-hundred-million-yuan Pre-A disclosed in May, only three months earlier.

Two rounds in three months is rare in the current funding climate. So is the company’s scale: founded two years ago, its daily token calls have already reached the trillions, and its customers and partners span internet firms, model vendors and leading enterprises across industries.

Capital is chasing a market whose barrier looks like it is falling. Open-source models from DeepSeek, Kimi, GLM and MiniMax ship dense releases, and firms can download weights directly while open inference frameworks mature. Some claim model inference is becoming a business “anyone with GPUs can run.”

The reality is different. “The models are open source, but the efficient deployment methods are not all open source,” Moxing AI founder and CEO Xu Lingjie told Zhidx. A model running at all only means tokens are produced. Whether those tokens can be sold at scale depends on holding latency, throughput, success rate and cost under heavy concurrent load.

Moxing AI says its daily sold tokens already reach the trillions, a figure drawn from tokens actually called and paid for by customers, not theoretical capacity after deployment. Revenue has reached hundreds of millions of yuan. Some of its models have already hit break-even, and the business is moving toward scaled profitability. Its next target is to push daily token output toward 10 trillion.

The business model serves large-B customers first. Major internet firms demand extreme stability and performance, and suppliers usually pass multiple test rounds before entering production. Once validated, customers expand share and migrate more models. Jin Chen, the CTO, notes that first-token latency above three seconds is already unacceptable to many enterprise clients. P50 and P99 latency, KV-cache hit rate, concurrent throughput and model precision all shape the choice.

Xu uses a telling example: with 80 desktop GPUs running a top model, single-stream speed can stall at just 20 tokens per second. The model runs, but the speed cannot meet commercial needs and the cost cannot support a competitive token price. In some scenarios, compute purchase and running cost can be 80 or even 90 per cent of total token cost.

Moxing AI sees three layers of moat. The first is the inference engine, optimising prefill-decode separation, memory management, load balancing and multi-hardware adaptation. The second is the supernode, which links more accelerators through high-bandwidth, low-latency interconnect to carry larger parallel compute, and human interaction feels smooth at 100 to 200 tokens per second, while complex agents may demand 1,000 tokens per second. The third is software-hardware co-design, already adapted to multiple domestic and overseas chips.

Xu frames the global AI contest as a shift from a model race to a deployment-capability race. As a basic production factor of the intelligent age, token production capacity is becoming a key measure of each nation’s AI industrial strength. Moxing AI wants to sit at the centre of both the token supply chain and the supernode industry chain as a full-stack AI infrastructure company.

Editor’s note: This is an adapted translation of the original Zhidx report. It has been trimmed and restructured for readability for an international business audience.

Leave a comment