OpenAI’s GPT-6 Astra rollout caps two days of launches that made model rankings obsolete in hours

Two days, three frontier launches, and the fastest overtake took five hours. On 2 September Google and Meta each released a new model: Gemini 3.8 Flash, Google’s third Flash iteration in six weeks, and Muse Spark 1.3, Meta’s fourth Spark in five months. Within hours Meta’s model had overtaken Google’s same-day release on the Artificial Analysis index. Forty-eight hours later OpenAI began rolling out GPT-6 Astra to all Plus subscribers, a staged launch it said would take days.

Frontier model release timeline chart showing OpenAI Google and Meta launches
Three frontier launches in two days compressed model lead times to hours. (Source: Sohu Tech)

Two routes collide on the same leaderboard

Google is playing a cadence game. Gemini 3.8 Flash prices at USD 0.75 per million input tokens and USD 3.75 for output, an entry price valid only until 31 December, and Google’s own tests show it closing on flagship models: 73.7 per cent on DeepSWE v1.1 against 74.0 per cent for Claude Opus 5, and 54.9 per cent on HLE-Verified against 54.4 per cent for Opus 5 and 54.5 per cent for GPT-5.6 Sol, at an output price roughly one seventh of Opus 5’s USD 25.

Meta is playing a summit game. Muse Spark 1.3 scored 62 on Artificial Analysis, behind only Claude Fable 5.1 at 66 and Claude Opus 5 at 63, up from 57 a month earlier and 53 in July. Meta chief executive Mark Zuckerberg called the model’s performance beyond imagination, and the company’s chief AI officer taunted Google on social media with “gemini who?”. For both companies the leaderboard lead was short-lived, because OpenAI’s GPT-6 Astra arrived 48 hours later.

When the measuring stick itself breaks

OpenAI’s launch is a scale play. GPT-6 Astra is rolling out to all Plus users rather than only Pro, Business and Enterprise tiers, and the company says it is a new system running at full production scale for the first time, mobilising large compute resources. But the scores carry a warning. ARC Prize analysis shows GPT-6 Astra at 62.7 per cent on ARC-AGI-3 under the standard framework, jumping to 99.9 per cent under a new Provider Adapter framework, a 37-point swing on the same model.

Across the board, frontier labs are converging. Gemini 3.8 Flash, Claude Opus 5 and GPT-5.6 Sol sit at 54.9, 54.4 and 54.5 per cent on HLE-Verified, separated by decimal points. Google admits its Flash model works harder, running more reasoning steps and tool calls that consume more tokens. When gains come from added compute rather than paradigm shifts, the frontier race becomes an exercise in diminishing returns, and the word “top” loses its meaning as the ruler itself wobbles.

Cost per task becomes the new metric

With intelligence converging, the decisive number is the cost of a unit of capability. Artificial Analysis puts Muse Spark 1.3 xhigh at USD 0.55 per task against USD 0.94 for Grok 4.6, USD 0.95 for GPT-5.6 Sol and USD 1.23 for Claude Opus 5 at similar scores. Gemini 3.8 Flash sits at USD 0.58. DeepSeek earlier cut V4-Pro prices permanently by 75 per cent, and Google is now using a dated entry price to lock in customers, evidence that the high-price closed model is losing its footing.

Google’s least-noticed move may be the most significant: Gemini 3.8 Flash Cyber, a cybersecurity-specialised model available only through its Fairwind programme to trusted government agencies, critical-infrastructure operators and software maintainers. Google claims 71.0 per cent on vulnerability discovery and 47.2 per cent pass at one on CWE-Bench patch repair, near Claude Fable 5’s 47.8 per cent. Security, the report argues, is becoming the least commoditised business in an age of commoditised intelligence, a market with an admission gate rather than an open API.

Who wins when everyone converges

The industry’s old certainties have inverted. The most expensive lab, Anthropic, stays closed and premium. The cheapest Chinese players keep cutting prices. The former open-source standard-bearer, Meta, keeps its strongest flagship closed. Google keeps releasing versions while digging a moat with a gated security model. Nobody is on a pure path anymore.

For developers the practical advice is blunt: stop paying a premium for the newest and strongest, buy on cost per task and success rate, and watch for models that inflate scores with extra reasoning steps while quietly raising token bills. The race is no longer about who is smartest. Intelligence is becoming a utility, and what is actually valuable is the gate, the pipe and the access card that decides who switches the light on.

Editor’s note: This is an adapted translation of the original Sohu Tech report. It has been trimmed and restructured for readability for an international business audience. The full original (in Chinese) is at https://www.sohu.com/a/1071692687_116132.

Leave a comment