AI · Opinion

Gemini 4 Argon's most important number is its 95% cache discount

Argon's benchmark wins are mostly Google grading Google. The pricing matters more: cheap re-reads make cost per finished task the way to pick your default model.

The most important number in Google’s Gemini 4 Argon announcement sits in one sentence about pricing. It is 95%, the discount on cached input tokens. At an introductory $2 per million input tokens, a cached token costs ten cents per million. Output is $10 per million. Almost everything else in the post is a ranking, from 77.9% on DeepSWE v1.1 to first place on the Vals Index. I think rankings are the wrong way for a startup to choose its default model, and Argon’s price sheet shows why.

I expected to spend this column picking holes in the benchmarks, and the holes are there. DeepSWE v1.1 is new to me, and I couldn’t find anyone outside Google reporting scores on it. The Vals Index weights finance, coding, legal and tax work by each sector’s share of US GDP, so a top spot there partly reflects somebody’s weighting choice. The 51.3% on Zapier’s AutomationBench and the 91.7% on LVBench are Google’s own reported figures. The internal anecdotes are vivid: more than 300 TiB of memory freed across Google’s data centres, and 32K lines of SIMD code in the libgav1 video decoder replaced with safe Rust that runs 2.7x faster than the earlier Rust port. Google picked these anecdotes itself, and I found no outside check of any of them.

What surprised me is how little the benchmark argument matters once you look at the bill. A coding agent spends most of its money re-reading things. The same repository, the same system prompt and the same tool definitions go back into the context on every turn. With a 95% cache discount those re-reads cost almost nothing, so output dominates the bill. Argon’s $10 per million output tokens matches the launch figures I have for Gemini 2.5 Pro and GPT-5, and it is below the $15 Anthropic listed for Claude Sonnet 4.5 and the $25 for Opus 4.5. I couldn’t re-check those list prices this month, so treat them as leads. But if output prices are roughly level and repeated input is nearly free, the cheapest model is the one that finishes the job in the fewest output tokens and the fewest attempts.

Argon’s other headline feature cuts both ways here. Google has raised the output limit to 1M tokens, up from 64K, and says the extra room lets the model crack hard problems “in one go”. A single trajectory that uses the whole allowance costs $10 in output alone, and hidden reasoning tokens are billed as output. Independent testers such as Artificial Analysis have often found that reasoning models cost several times more per task than their list prices suggest. A model that thinks for a long time and gets the answer right once can work out cheaper than a terse model that needs several retries and an engineer to tidy up after it. It can also spend ten dollars on a task a smaller model would have finished for pennies. Google’s post won’t tell you which, because it publishes scores and never tokens per task.

The obvious reply is that none of these numbers deserve trust. METR’s randomised trial in July 2025 found experienced open-source developers took 19% longer with AI tools, while believing they had been about 20% faster. The April 2025 paper “The Leaderboard Illusion” reported that big labs privately tested many model variants on LMArena and published only the best. I accept both findings, and I think they support my argument. Cost per accepted unit of work, measured on your own tasks, doesn’t depend on Google’s charts being honest. The METR result tells you to measure, and it gives no reason to stay with whichever model you picked last year. It also means Argon’s real rival for bulk work may be Google’s own Gemini 2.5 Flash, which listed at roughly $0.30 input and $2.50 output per million.

Nobody outside the Fairwind Program can use Argon yet. Google is giving it to vetted cyber defenders first, and the $2 and $10 prices are labelled introductory, so the sums could change before paid API customers and Google AI Ultra subscribers get access. Even so, I expect that once it opens up, agent-heavy startups with high cache-hit rates will find it the cheapest capable option per finished task. I expect that to hold even if the DeepSWE score drops once independent testers run it.

You can prepare now. Export last month’s agent logs and split every call into uncached input, cached input and output tokens. Price each bucket at Argon’s three rates, then divide the total by the number of tasks that actually merged or shipped. Do the same at your current model’s rates. When Argon reaches the paid API, rerun a sample of those same tasks through it. If the cost per shipped task is lower and the retry rate is no worse, change your default model that week and stop checking leaderboards.

Prompted by Gemini 4 Argon: our next era of frontier intelligence, Google.