DeepSeek V4 Flash is redrawing how models are judged. On OpenCode, a single AI coding tool, the model burned 8 trillion tokens in one day, more than the entire OpenRouter platform’s daily average of about 6.6 trillion across 400-plus models.
The reason is price. DeepSeek charges about 2 yuan per million output tokens; Anthropic’s Claude Opus 4.8 is priced at 25 dollars, roughly 170 yuan, an 85-fold gap. On Artificial Analysis’s Intelligence Index, V4 Flash Max scored 50 to Opus 4.8 Max’s 56, yet ran the full eval for about 72 dollars versus 3,752 dollars: 52 times cheaper.
DeepSeek did not top the charts. It changed the unit of comparison. In the agent era a task loops for hours, calling tools dozens of times; token use can balloon a hundredfold versus a single answer. So the question shifts from ‘how smart per reply’ to ‘how much does a long task cost, and how often does it fail.’
That is the ‘kill line’: a model need not be the strongest, only ‘good enough most of the time,’ to use a two-orders-of-magnitude price gap to push pricier models out of the default slot. The awkward middle, stronger than V4 Flash but costing tens of times more, is what gets squeezed.
Architecture explains the cost: V4 Flash has 284 billion parameters but activates only about 13 billion per token via DeepSeekMoE and compressed attention (CSA/HCA), with KV cache at about 7 per cent of V3.2. Post-training zeroed in on coding and agent-task failures. And it now runs on a high-memory workstation (DwarfStar’s quantised build fits 128GB Macs), chipping away at the ‘frontier models live only in data centres’ wall.
Original report: Sohu IT.
Translated and adapted from Sohu IT (it.sohu.com).