Google announced Gemini 3.5 Flash at I/O 2026, and the pitch was speed and price. Sundar Pichai said it runs at 289 tokens per second, four times faster than other frontier models, at less than half the cost of comparable frontier models. This is true, and it is also a specific kind of true. The comparison is to other companies’ flagship models. The comparison Google did not lead with is to Google’s own previous Flash.

Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens. Gemini 3 Flash, the model it replaces, cost $0.50 and $3.00. That is exactly three times more, on both ends. Gemini 2.5 Flash cost $0.30 and $2.50. The original 1.5 Flash billed output at roughly thirty cents per million tokens, which makes the new model around thirty times more expensive to generate text than the model that established what the word Flash was supposed to mean.

Flash was supposed to mean cheap. It was the small fast tier, the one a developer reached for when a task did not justify a flagship. Gemini 3.5 Flash is priced like a flagship and named like a bargain. Google’s own framing gives this away: the model is cheap relative to GPT-5.5 and Claude Opus 4.7, not cheap relative to last quarter. “Affordable” has quietly become a claim about competitors rather than a claim about Flash.

There is a reason the price moved, and it is the more interesting part. On the benchmarks Google chose to show, Gemini 3.5 Flash is genuinely strong, and it is strong specifically at agentic work. It scores 76.2 percent on Terminal-Bench 2.1, up from 70.3 for Gemini 3.1 Pro. It posts 83.6 on MCP Atlas and 1656 Elo on GDPval-AA, beating Google’s own larger model. A Flash that beats the Pro is an excellent thing to ship and a slightly awkward thing to headline a keynote with, because it raises the question of where the Pro is.

It also raises the question of what was traded away. On pure reasoning, the new Flash moves backward against Google’s previous model: 40.2 percent on Humanity’s Last Exam versus 44.4 for Gemini 3.1 Pro, with a lower ARC-AGI-2 as well. The model was tuned to act rather than to reason at length, and it was repriced to match. The cheap tier got more expensive because making it competitive on agentic tasks was not free. The price is the receipt for the capability, and the receipt says the gap is not closing on its own.

Closing distance to the frontier is getting more expensive, and the people who build the frontier are concentrating. On the same day as the Flash launch, Andrej Karpathy, an OpenAI co-founder and the former head of Tesla’s self-driving program, announced he had joined Anthropic. He is not joining to work on products. He is starting a team focused on using Claude to accelerate pre-training research, under pre-training lead Nick Joseph. Pre-training is the part of the stack that produces raw capability in the first place.

None of this makes Gemini 3.5 Flash a bad model. It is a good model. It is fast, it is strong at agentic work, and it beats Google’s own larger model on the benchmarks Google picked. The one thing it is no longer is cheap, which is the thing Flash was for. So Google built a very good model, charged Pro money for it, and kept the Flash name, and on the same day Andrej Karpathy went to Anthropic to work on pre-training. Catching up, it turns out, is expensive, and the people who are best at not having to catch up mostly work at the other two labs.

–
By the Control Plane Editorial Team