Google announces Gemini 4 Argon, skipping the promised 3.5 Pro — and you still cannot use it
Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.
Published entries across all sections carrying the “Model inference” tag, newest first by publication date on this site.
5 entries
Google announced Gemini 4 Argon on September 30: 77.9% on DeepSWE above GPT-6 Astra, a 1M-token output limit, phased release, no API pricing yet.
DeepSeek published six Ascend component repos on September 30; it says every V4-series NVIDIA operator now has an Ascend counterpart.
OpenAI will run managed agents on AWS and bring ChatGPT to Slack and Teams without individual licenses, with a private-inference preview coming this fall.
Inferact, founded by the original vLLM team, wrote a megakernel inference kernel for Google TPUs. Paired with DeepSeek’s DSpark speculative decoding, 16 TPU v7 chips served Kimi K3 at 709 tokens per second versus 452 on GB200 under the same setup, QbitAI reports. The code is open source.
Free online chat access and WebGPU-accelerated in-browser local inference for the open-weights DeepSeek R1 reasoning model.