Gemini 4 Argon: Google Is Back in the Top 3 — and It's Cheaper Per Task

Gemini 4 Argon intelligence index chart

Google just stopped being a footnote.

Gemini 4 Argon — DeepMind's first proprietary model above the Flash class in over seven months — scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max) and edging GPT-6.1 Sol (max) by a point. That puts Google back in the top three labs on raw intelligence for the first time in a while.

But the number that matters for anyone running agents is the cost per task.

The agent economics

At its current 50% launch discount, Gemini 4 Argon runs $1.99 per Intelligence Index task — about 60% of what GPT-6 Astra (max) costs for comparable intelligence. Cache reads get a 95% discount, up from 90% on Gemini 3.8 Flash. When the promo ends, that climbs to $3.98, so the window to test it cheap is now.

Gemini 4 Argon cost per task chart

It's not the cheapest frontier option — GPT-6.1 Sol still undercuts it at $0.72 per task. But Argon buys something Sol doesn't: agentic reliability.

Why this matters for your agents

Gemini has historically been the weak link on agentic work. Argon flips that:

➤ #1 on AutomationBench-AA at 78% — 7 points ahead of Claude Sonnet 5.5 (max). This is the benchmark that measures whether an agent can actually finish a real automation task.

➤ Terminal-Bench 4.0 at 57% — a +53 point jump from Gemini 3.1 Pro Preview, now only behind Claude Sonnet 5.5, Opus 5.5, and GPT-6 Astra.

➤ Lowest hallucination rate among leading models — 15% on AA-Omniscience, versus 51% for GPT-6 Astra and 54% for GPT-6.1 Sol. For autonomous agents that run without a human watching every step, this is the stat that saves you from silent, expensive mistakes.

Gemini 4 Argon agentic performance chart

What to actually do

1. Test it during the discount. At $1.99/task with 95% cache discounts, the economics are unusually favorable for agent workloads that resend long context each turn.

2. Watch the token profile. Argon averages 62k output tokens per task vs 27k for GPT-6 Astra — cheaper per token, but it talks more. Run your own cost-per-completed-task measurement, not just per-token price.

3. Use it where hallucination hurts most. If your agent writes to a database, sends emails, or touches money, the 15% hallucination rate is the strongest argument for Argon over the GPT-6 family.

Gemini 4 Argon is rolling out to selected users and isn't publicly available yet. The 50% discount is a launch promotion with no confirmed end date. If you run agents, this is the model to benchmark against while the price is low.

Source: Artificial Analysis — "Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved" (Sept 30, 2026). Charts via Artificial Analysis.