Models#Benchmark#Model inference

Google announces Gemini 4 Argon, skipping the promised 3.5 Pro — and you still cannot use it

Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.

Ars Technica art: the Gemini sparkle on a black background

Google announced its Gemini 4 Argon flagship on September 30, jumping straight past the previously promised Gemini 3.5 Pro. The official numbers: 77.9% on the DeepSWE v1.1 coding benchmark, above GPT-6 Astra, Fable 5.1 and Opus 5.5, a 1M-token output limit, and chain-of-thought monitoring with stop systems. It is in a phased release — the general public cannot use it yet.

Facts

  • Benchmarks: 77.9% on DeepSWE v1.1 per Google, above GPT-6 Astra, Fable 5.1 and Opus 5.5 (all official comparisons).
  • Output limit: 1M tokens per response, up from 64K — aimed at long-generation tasks.
  • Safety: chain-of-thought monitoring and stop systems, alongside a DeepMind “reasoning transparency” essay.
  • Release cadence: “models of this scale call for a phased release” — trusted testers first (including Fairwind cybersecurity partners such as Wiz), then paid API plus AI Ultra, then enterprise, then consumers; no API pricing, no firm GA date.
  • Internal claims: fleet-telemetry optimizations across 300 TiB of memory; 800K+ lines of C/C++ moved to Rust in the Fuchsia Zircon kernel.

Editorial take

Skipping version numbers and shipping-while-unavailable is the new flagship rhythm — OpenAI halted GPT-6.1 outright, Google pre-announces scores and trickles out access; different routes, same lesson: version numbers stopped being a capability calendar. Hold procurement until API pricing lands, and treat DeepSWE as a Google-chosen benchmark — wait for third-party re-tests before comparing against GPT-6.1 Sol’s pricing.