Google announces Gemini 4 Argon, skipping the promised 3.5 Pro — and you still cannot use it
Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.
Published entries across all sections carrying the “Benchmark” tag, newest first by publication date on this site.
8 entries
Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.
OpenAI shipped GPT-6.1 Sol at DevDay: near-Astra performance at one-fifth of Astra’s prices ($2/$10 per million tokens), live in ChatGPT Work and Codex.
Anthropic released Sonnet 5.5 on September 28: $2/$10 per million tokens (half of Opus 5.5), 30% faster, and better agentic coding than Opus 5.5 in-house.
Open-source SparkDiffusion pushes 720P-14B video generation to 265x on one RTX 5090 via sparse attention, distillation and FP8; VBench-2.0 drops 0.5 points.
Stealth model Pixel Canary went free on Vercel AI Gateway on September 25, tying GPT 6 Astra on Next.js coding evals — at nearly 17 minutes per task.
Shanghai Innovation Institute released the Nex-N2.5 family on September 9, 2026 — 35B, 397B and 1.6T models, all under Apache-2.0. On the official table the 1.6T Max tops BrowseComp at 92.6, ahead of Claude Opus 5 and GPT-5.6 Sol, and outscores GLM-5.3 and DeepSeek-V4-Pro on automation and tool-use benchmarks.
Anthropic released Opus 5.5, the first Claude 5.5 model, on September 23. The company says it matches Fable 5.1 on most tasks and costs 40% less per typical task than Opus 5, at $4 input and $20 output per million tokens.
Inferact, founded by the original vLLM team, wrote a megakernel inference kernel for Google TPUs. Paired with DeepSeek’s DSpark speculative decoding, 16 TPU v7 chips served Kimi K3 at 709 tokens per second versus 452 on GB200 under the same setup, QbitAI reports. The code is open source.