Models#Benchmark#Model inference
Google announces Gemini 4 Argon, skipping the promised 3.5 Pro — and you still cannot use it
Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.

Google announced its Gemini 4 Argon flagship on September 30, jumping straight past the previously promised Gemini 3.5 Pro. The official numbers: 77.9% on the DeepSWE v1.1 coding benchmark, above GPT-6 Astra, Fable 5.1 and Opus 5.5, a 1M-token output limit, and chain-of-thought monitoring with stop systems. It is in a phased release — the general public cannot use it yet.
Facts
- Benchmarks: 77.9% on DeepSWE v1.1 per Google, above GPT-6 Astra, Fable 5.1 and Opus 5.5 (all official comparisons).
- Output limit: 1M tokens per response, up from 64K — aimed at long-generation tasks.
- Safety: chain-of-thought monitoring and stop systems, alongside a DeepMind “reasoning transparency” essay.
- Release cadence: “models of this scale call for a phased release” — trusted testers first (including Fairwind cybersecurity partners such as Wiz), then paid API plus AI Ultra, then enterprise, then consumers; no API pricing, no firm GA date.
- Internal claims: fleet-telemetry optimizations across 300 TiB of memory; 800K+ lines of C/C++ moved to Rust in the Fuchsia Zircon kernel.
Editorial take
Skipping version numbers and shipping-while-unavailable is the new flagship rhythm — OpenAI halted GPT-6.1 outright, Google pre-announces scores and trickles out access; different routes, same lesson: version numbers stopped being a capability calendar. Hold procurement until API pricing lands, and treat DeepSWE as a Google-chosen benchmark — wait for third-party re-tests before comparing against GPT-6.1 Sol’s pricing.