Back to glossary

Term

Gemini 3 Flash

Gemini 3 Flash is Google's fast, low-cost thinking model from December 17, 2025 — near-Pro reasoning with a 1M-token context, built for agentic workflows and coding.

Gemini 3 Flash — explained in more detail

Google released Gemini 3 Flash on December 17, 2025 as the fast, cost-efficient variant of the Gemini 3 generation. Within the Gemini family, Flash is the efficiency tier: it sits below the flagship Gemini 3 Pro but targets high throughput at a low price. Gemini 3 Flash is a thinking model — it can run explicit reasoning steps before answering — and is designed for agentic workflows, multi-turn chat and coding assistance. Google highlights improved visual and spatial reasoning as well as agentic coding.

According to Google, Gemini 3 Flash delivers near-Pro reasoning roughly three times faster than Gemini 2.5 Pro while using about 30% fewer tokens.

Example / Practical use

The key data: a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. Pricing is $0.50 per 1M input tokens and $3 per 1M output tokens; cache reads and image input are cheaper, while web search is billed separately. Access is proprietary via the Gemini API and Google’s AI platforms. The large context window and low token price make Flash particularly suited to high-volume tasks — processing long documents, RAG pipelines, or agents that need many model calls per run.

Distinction

Gemini 3 Flash is not Google’s most capable model — for maximum quality there is Gemini 3 Pro. Flash trades some peak capability for speed and cost, making it the workhorse for scaling production loads. Within the Flash line, the 3 version is the successor to the 2.5 Flash generation.

See everything in one place:Gemini Flash