DeepSeek, once nicknamed the large-model “price butcher”, is sheathing the knife. At 00:00 on 17 August, its new API prices took effect.

The adjustment covers V4 Pro and V4 Flash and introduces peak-valley pricing for the V4 series. Beijing time 09:00 to 12:00 and 14:00 to 18:00 are peak windows; the other 17 hours are off-peak, priced at half the peak rate. After the change, flagship V4 Pro at peak charges 0.3 yuan, 9 yuan and 27 yuan per million tokens for cache-hit input, cache-miss input and output, rises of 1,100 per cent, 200 per cent and 350 per cent. V4 Flash peak prices also rose 400 per cent, 200 per cent and 350 per cent.
A price hike that was telegraphed
This was a pre-announced increase. As early as late June, DeepSeek told users it planned peak-valley pricing at V4’s official launch, to allocate resources reasonably and improve stability. The original plan put V4 Pro off-peak and peak at 0.025, 3 and 6 yuan, and 0.05, 6 and 12 yuan; V4 Flash peak was 0.04, 2 and 4 yuan. It was meant to land with the mid-July V4 launch but the model slipped.
On 6 August DeepSeek posted another notice and emailed some API users, saying it would raise API prices significantly and advising them to plan usage, without giving specifics. About a week later V4 Pro and the final pricing went live. V4 Pro targets complex reasoning, coding and agent tasks, with a 1 million token context, up to 384,000 token output, thinking mode, tool use and OpenAI Responses API support.
Four days later the new prices took effect. The final pricing sits clearly above the late-June plan. Take V4 Pro output: the plan’s off-peak and peak were 6 and 12 yuan, while the final off-peak and peak reached 13.5 and 27 yuan, with the off-peak price already above the earlier peak.
So this is not only peak-valley steering. Within the time-of-day frame, DeepSeek has raised V4 API prices across the board.
The gap with rivals narrows
At peak, V4 Pro cache-miss input and output are 9 and 27 yuan per million tokens. For comparison, Zhipu’s GLM-5.2 runs about 8 and 28 yuan, already close; Moonshot’s Kimi K3 is 20 and 100 yuan, still clearly higher. Models differ in ability and use case, so token price alone is not a verdict. DeepSeek’s edge has not fully vanished: V4 Pro off-peak cache-miss input and output are 4.5 and 13.5 yuan, still below those two, and its cache-hit input tops out at 0.3 yuan, far under GLM and Kimi at about 2 yuan.
After the hike DeepSeek is no longer an order of magnitude cheaper than some domestic flagships. Developers will shift from scanning price tables to counting calls, token burn and final result.
Who gets hurt
Long-running agent developers and teams are hit hardest. They burn large token volumes and keep long contexts, so they cannot easily cut cost by shortening chats. One AI blogger using V4 Flash for browser search, debugging and server work showed an 8 August log of 267 requests and about 76.73 million tokens, with 75.47 million cache-hit input, 1.05 million cache-miss input and 210,000 output. At old prices that cost about 2.99 yuan; on the new peak-valley schedule it rises to 7.27 yuan, 2.43 times more, and he called that light to moderate use. The cache-hit input rise, up to 1,100 per cent at peak for V4 Pro and 400 per cent for Flash, is the sharpest line item for his profile.
Which increase bites most depends on the token structure. High-cache, long-context agent users feel the cache rise; output-heavy work like marketing copy, product blurbs and long reports feels the output rise, where V4 Pro peak output jumped from 6 to 27 yuan per million tokens.
Editor’s note: This is an adapted translation of the original Sohu Tech report. It has been trimmed and restructured for readability for an international business audience.