Skip to content

The Fast tier

The Fast tier runs gemini-3.6-flash by default — $1.80 per million input tokens and $9.00 per million output tokens. What lands here, what it costs, and what happens when it cannot run.

Fast is one of four options in the composer’s engine menu, alongside Auto, and the other two tiers. It exists to give you the cheapest acceptable answer.

The Fast tier runs google/gemini-3.6-flash by default, made by Google.

Runs as Input / 1M tokens Output / 1M tokens
Fast tier default $1.80 $9.00
Fast tier default, repeated context $0.18

The second row is the repeated-context rate. When a request re-sends context the model has already seen, those tokens are billed at that lower rate instead of the full input rate — it is applied automatically by the biller, not something you switch on. See How billing works.

You can point this tier at a different model in Settings → How it worksSet an engine per speed tier, and you can override it for a single request from Pick an exact engine.

Send work here when the answer is short, the material is small, and being right matters more than being thorough: rewriting a paragraph, extracting a few fields from a page you paste in, drafting a reply, classifying or tidying a list.

Work that does not belong here is anything with many steps that depend on each other, or anything where you would rather pay more than re-do it. That is what Max is for.

Two things this page deliberately does not tell you: how Auto decides, and how a request moves between models. Those are internal and we do not publish them. What we do publish is the above — pick a tier and you know which model runs and what it costs.

Both examples below run as the Fast tier default, at the two rates above:

Runs as Run shape Cost
Fast tier default 4,000 in, 800 out (a short question with a little context) $0.0144
Fast tier default 100,000 in, 10,000 out (a long document and a long answer) $0.27

That is the whole arithmetic: tokens × the two rates. Nothing is rounded up, there is no per-run fee, and the receipt itemises every line. The ceiling on one request is this model’s context window, 1M tokens (about 786,000 words, roughly 1.6k pages) — instructions, attachments and the answer all share it.

If gemini-3.6-flash loses its price or its capacity, this tier does not silently run something else:

  • If you had re-bound Fast to your own choice and that choice lost supply, the tier falls back to its own default and tells you it did — you get a notice naming what you asked for and what actually ran.
  • If the default itself is unavailable too, the tier reports itself unresolved rather than borrowing another tier’s model. There is no quiet cross-tier substitution.

This is the only failure behaviour we publish, because it is the only one we can point at in the code that implements it.

You will notice: which tier you are on (the engine menu shows it), a fallback when one happens (it is reported, not hidden), and every charge (itemised on the receipt — see Reading a receipt).

You will not notice, because we do not expose it: how Auto chooses, or anything about where capacity comes from.

  • We publish no measured speed for this tier. Tiers are named for what they are for, not for how fast they run: “Fast” is a meaning, not a benchmarked latency, and we have no timings we would stand behind.
  • We publish no measured quality comparison between tiers. Nothing here claims Fast answers better or worse than another tier on your work.
  • Image input on gemini-3.6-flash has not been measured. If you need to hand over a picture, use a model from Models that can read your images instead of assuming this tier can.

Prices last changed 2026-07-29 (source ced6a4f). Every figure here is generated from the same pricing code that bills your runs.