DeepSeek raises API prices, ends China LLM war

A week after its cheapest-ever model sparked talk of a pricing kill line, DeepSeek announced it will raise API prices. On 6 August the company told its developer platform that pricing would rise soon by a large margin, and users began receiving emails. The message was blunt: keep using after the change and you accept the new rates, or exit and request a refund.

DeepSeek model API pricing dashboard
DeepSeek will raise API prices as call volume strains its compute. (Sohu)

For a brand built on being cheap, the reversal is striking. Two years ago its V2 model priced input at 1 yuan and output at 2 per million tokens, about one hundredth of GPT-4-Turbo and the lowest then on the market. That earned the price-butcher name, and its open-source stance earned the source-god nickname.

V4-Flash-0731, released a week earlier, set an all-time low and triggered a global kill-line debate. Against Kimi K3, Qwen, GLM-5.2, GPT-5.6 and Claude, its performance gap is modest. Artificial Analysis ranks it level with GPT-5.6-Luna and Gemini-3.6 Flash, behind the top Chinese and US models. But it is the lowest-cost model in the world, about 0.03 US dollars per task on average, under 4 per cent of Kimi K3 and under 1 per cent of Claude Fable 5.

The developer community pushed back hard, and Liang Wenfeng’s reputation slid from Saint Liang to Liang the liability. Most developers read the hike not as greed but as surging call volume straining compute, in the words of one, the boss says he does not want this money, they are forcing him to earn it.

The load is real. OpenCode says V4-Flash once processed 8 trillion tokens in a day, and its users spend 130,000 US dollars daily on DeepSeek, with frequent outages. On OpenRouter, V4-Flash topped the global weekly chart at 7.22 trillion tokens and held 5.71 trillion this week, with Chinese models filling the top five.

DeepSeek is not first to move. Zhippu raised GLM-5 pricing through the year, with its domestic coding plan up 130 to 260 per cent. Kimi K3 output pricing rose 2.7 times over its prior model. Alibaba, Baidu and Tencent cloud lifted AI compute and storage 5 to over 30 per cent. Jefferies puts US LLM API prices up 89 per cent in the second quarter, with Claude Fable 5 doubling. A few cut instead, Xiaomi permanently dropped MiMo prices up to 99 per cent, and OpenAI cut GPT-5.6 Luna by 80 per cent.

The strategic read is that China’s model makers are trying to shed the low-price tag and claim pricing power through capability, but scarce compute is the real driver. Goldman Sachs lifted its estimate of combined China LLM annual recurring revenue to 13 billion US dollars by year end and 125 billion by 2030, a 25-fold rise in five years, though the sector stays loss-making until about 2030.

DeepSeek itself is a variable. Annualised revenue nears 500 million US dollars with flagship API gross margin of 70 to 80 per cent, well above peers. It closed a first 50 billion yuan round and is seeking a second near 8 billion US dollars, partly for compute and data centres. Once supply catches up, a cut is possible again. For now, the paused price war may simply be waiting for the next storm.

Editor’s note: This is a translated adaptation of a Chinese-language report from Sohu Tech (sohu.com). Figures, dates and direct quotations are reproduced as published.

Leave a comment