Back to glossary

Term

Gemini 3.6 Flash

Gemini 3.6 Flash (July 2026) is Googles fast, cost-efficient Flash tier with a 1M-token context — built for high throughput and low cost per request.

Gemini 3.6 Flash — explained in more detail

Gemini 3.6 Flash was released by Google on 21 July 2026 as a production-ready model. Within the Gemini line, Flash is the fast, cost-efficient tier: not the strongest model (that is Gemini Pro) but the one built for high throughput, low latency and low cost per request. Flash is typically used wherever large volumes of requests need to be answered quickly and cheaply.

Like the rest of the Gemini line, the model is natively multimodal (text, images, audio, video) and keeps the characteristic context window of around 1 million tokens (exactly 1,048,576), with up to 65,536 output tokens per response. Despite its positioning as the cheap tier, 3.6 Flash improved markedly over its predecessor 3.5 Flash: SWE-Bench Pro 58.7 percent (up from 55.1), DeepSWE 49 percent (up from 37), MLE-Bench 63.9 percent (up from 49.7) and, on the long-context test GDM-MRCR v2 over 1M tokens, 54.0 percent (up from 26.6). On throughput the model reached around 303 tokens per second — several times the median of comparable models. On price it sits at roughly 0.75 US dollars per million input and 3.75 US dollars per million output tokens.

Access is proprietary: Gemini 3.6 Flash runs through the Gemini API, Google AI Studio and Vertex AI. The weights are not public.

Example / Practical context

In practice, Flash is the default choice for high-volume, tight-budget tasks: classification, summarisation, extraction, chat responses, or as a fast first stage in a pipeline that hands hard cases to a stronger model. A typical setup: a service processes tens of thousands of incoming documents a day, lets Flash categorise and summarise them, and escalates only the uncertain cases to a Pro model. The high throughput and low token cost are exactly what make this pattern economical.

Within the Gemini line, Flash sits between the smaller Flash-Lite variants (cheaper, weaker) and the Pro flagship (stronger, pricier, slower). The 3.6 designation marks an intermediate step within the third Gemini generation; later versions (such as Gemini 3.7 Flash) supersede it. Open-weight models (such as Llama, Qwen or DeepSeek) differ fundamentally from the API-only Gemini models: their weights can be downloaded and run locally, whereas Gemini is available solely as a hosted service. Gemini the model should not be confused with the former Bard branding or with the consumer app of the same name.

See everything in one place:Gemini Flash