On 31 July, OpenAI cut API prices across the GPT-5.6 family. The steepest cut hits Luna, down 80 per cent, input to $0.2 per million tokens, output to $1.2. The balanced Terra model drops 20 per cent, input $2, output $12 per million. The flagship Sol holds its price but gains a Fast mode running up to 2.5x the standard inference speed at double the cost.
The move follows a self-improvement milestone. A day earlier OpenAI disclosed that GPT-5.6 Sol had been used to optimise its own inference system, autonomously rewriting production GPU kernels, running hundreds of experiments and cutting overall serving cost by 20 per cent while lifting token-generation efficiency more than 15 per cent. The price cut is that efficiency, handed back to developers.
On Artificial Analysis’ price-performance map, Luna sits in the top-left, 50-plus intelligence at roughly $0.05 a task, undercutting Claude Opus 5, Gemini 3.6 Flash and GLM-5.2 Max at similar scores, and richer than DeepSeek V4 Pro. OpenAI is framing Luna as the best cost-intelligence balance in the frontier.
The developer reaction was immediate. Some cheered the relief on their API bills, others tagged Anthropic with a your-move meme, and a few lined Luna up against Kimi K3 on price. Altman’s line, the best price-to-intelligence tradeoff at every tier, signals where the next model war is fought: not just smarter, but cheaper per unit of thought.
Read the original report (Sohu IT)
Translated and adapted from Sohu IT (it.sohu.com).