News

The Gemini 3.8 Flash price has an expiry date

Google released Gemini 3.8 Flash on September 2 at $0.75 per million input tokens and $3.75 per million output tokens, the same headline price as 3.7 Flash three weeks earlier. The number worth reading is in the footnote: that price is introductory and expires December 31, 2026. On January 1 it becomes $1.50 and $7.50.

A doubling with a date on it

Most price changes arrive as a surprise email. This one arrived with the launch post, four months ahead, in a footnote under the fold. Google is telling anyone building on 3.8 Flash exactly when their inference bill doubles, which is more warning than the industry usually gives.

It also reframes what the launch price means. Google has now shipped three Flash releases in six weeks, each at $0.75 input. Read the footnote and the pattern looks less like cheap inference and more like a trial period that the whole Flash line has been running inside.

The model is built to spend more tokens

Price per token is only half of a bill. The other half is how many tokens the model decides to use, and Google is unusually direct about this one using more. The launch post says 3.8 Flash "works harder" on complex tasks, runs extra reasoning steps, calls tools repeatedly, and may burn more tokens at higher effort levels. Output pricing includes thinking tokens.

You do get something for it. On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash scores 73.7% against 65.3% for 3.7 Flash, which puts it within a rounding error of Claude Opus 5 at 74.0%. It also reports 54.9% on HLE-Verified. Those are Google's published figures.

So the honest version of the pricing question is not what a million tokens cost. It is what one finished task costs, and that answer moved in two directions at once. Google left 3.7 Flash fully supported and recommends it when compute efficiency matters more than the last few benchmark points, which is a quiet admission that the newer model is not always the cheaper one.

What we do about it

We log tokens per completed task, not tokens per call. A model that answers in one pass at higher effort can beat a cheaper model that needs four attempts, and neither the price sheet nor the benchmark table will tell you which one you have. Only your own traces will.

Effort level is a config value in our projects, not a constant buried in a client wrapper. Most product work does not need maximum reasoning to fill a form or classify a support ticket, and being able to dial that down per feature is what turns a January price change into a settings adjustment rather than a budget meeting.

Sources

Have a stack outgrowing itself?

Book a call and we'll walk you through how we'd approach your platform, priced fairly and estimated for real.