China’s ‘Cheap’ AI Models Are Quietly the Expensive Ones

This past month Silicon Valley shipped GPT-5.6, Fable 5 and Grok-4.5, and China’s developer community lit up with benchmarks and debates. Yet for all the buzz, the country’s homegrown large models remain in a strange state: critically acclaimed, commercially cold.

The usage-versus-revenue gap is widening. OpenRouter data for 22-28 June shows Chinese models handled 20.39 trillion tokens a week against 4.25 trillion for US models – leading for nine consecutive weeks. But the financials tell a different story. Public filings put Zhipu’s 2025 revenue at RMB 724 million against a loss exceeding RMB 3 billion; MiniMax booked $790 million (about RMB 570 million); Moonshot and the rest have not disclosed full-year revenue. Even the leaders are stuck at the hundreds-of-millions level, while Anthropic’s annualised revenue has reached an estimated $60-70 billion, driven by B2B APIs and AI coding.

The counterintuitive truth is that Chinese models are, in a sense, expensive. On token-unit price they look cheap: GLM-5.2 lists at $1.4/$4.4 per million tokens in and out, DeepSeek-V4-Pro at $0.435/$0.87, against GPT-5.6 Sol at $5/$30 and Claude Fable 5 at $10/$50. But in high-intensity coding, where subscription bundles flatten cost, GPT and Claude are far cheaper.

One AI researcher who burns several billion tokens a day estimated that the same ~4 billion tokens would cost about $1,084 (roughly RMB 7,800) on GLM-5.2, versus almost nothing under OpenAI’s Codex plan – Plus at $20 a month, Pro from $100, with five to twenty times the allowance. Claude Code runs a similar structure: Pro at $20, Max 5x and 20x at $100 and $200. Under that cover, enormous token bills simply disappear.

For Chinese developers, burning tens of billions of tokens a day is routine. One heavy coder spent just RMB 300 a week on Codex – even without reimbursement – and still got more volume than GLM-5.2 would sell for the same money. Domestic plans tell the opposite story: Kimi Code tiers at RMB 49-699 refreshed weekly; GLM’s coding plan allows 80/400/1,600 prompts per five hours with triple deductions at peak; Alibaba’s Qoder enterprise plan can expire in days. In January 2026, amid a compute shortage, Zhipu cut its daily saleable volume to 20% and rationed allocations each morning – ‘harder to buy than concert tickets,’ as one report put it.

The root cause is compute scarcity. China operates 449 data centres against America’s 5,427, with total compute of 1,053 EFLOPS versus 2,400; much of China’s capacity rides Huawei Ascend and Cambricon chips a generation behind the H100. From March 2026, Zhipu, Alibaba Cloud, Tencent Cloud and Baidu raised token prices between 5% and 460%. The ‘low-price war’ is over.

The business logic is brutal. OpenAI and Anthropic sell to high-value enterprise – Cursor and GitHub Copilot alone contribute about $1.4 billion a year to Anthropic – turning fixed monthly fees into long-term lock-in with high ARPU. Chinese cloud, by contrast, is thin: Alibaba Cloud’s FY2026 adjusted EBITA margin was just 7.5%. And unlike a cloud server that sits at 10% utilisation and can be oversold ten times, every inference burns real GPU, HBM and electricity with no slack.

Performance is no longer the only wall. GLM-5.2 is now seen as the first domestic model to reach ‘Opus level’ on genuine complex, long-horizon tasks, though it still trails on world knowledge; caching inefficiency and peak throttling sour the experience. Kimi K3 – 2.8 trillion parameters, open-weight, a one-million-token context – scores 57 on Artificial Analysis (third globally) and tops the Arena frontend-coding board at 1,679, yet at RMB 2/20/100 per million cached/in/out it is no bargain against Codex, and developers report API dropouts.

The bottom line: Chinese developers are not unwilling to pay. The models simply cannot yet offer Codex-class value bundles because compute is tight and costly. The win waits on domestic compute at scale – and on price, stability and performance forming a combined front.

Source (Chinese original): Sohu IT

Translated and adapted from Sohu IT (it.sohu.com).

Leave a comment