Term
Gemini 3.5 Flash
Gemini 3.5 Flash is Googles high-efficiency multimodal model from May 2026 — it brings near-Pro coding and reasoning quality at Flash-typical cost and speed.
Gemini 3.5 Flash — explained in more detail
Gemini 3.5 Flash is Googles high-efficiency multimodal model and was released on May 19, 2026. Googles Flash line stands for the balance of speed, cost and quality; with version 3.5 it moves markedly closer to the Pro class, delivering — per Google — near-Pro coding and reasoning performance at Flash-typical latency. The model accepts text, images, video, audio and PDF as input and produces text output; tool calling and structured JSON outputs are supported.
Key facts
- Release: May 19, 2026, proprietary access via the Gemini API, Google AI Studio and Vertex AI.
- Pricing: $1.50 per 1M input tokens, $9 per 1M output tokens; cache read $0.15, audio input $3 per 1M.
- Context window: 1,048,576 tokens (1M), maximum 65,536 output tokens.
- Knowledge cutoff: January 1, 2025.
- Benchmarks: GPQA Diamond 92.4 percent, TAU-Bench 75.3 percent.
Example / Practical use
Flash 3.5 is optimized for high throughput and parallel agent execution — at roughly 189 tokens per second and a P50 latency of about 0.79 seconds it suits chat, extraction, summarization and coding assistance at large volume. Its multimodality additionally makes it usable for image, video and document processing without deploying a separate model per modality.
Delimitation
Within the Gemini 3.5 line, Flash sits between the cheaper, leaner Flash Lite variant and the more capable Pro class. Where maximum reasoning depth matters, Gemini Pro remains the choice; for latency- and cost-critical bulk work, Flash is the default. Direct competitors in the same efficiency class are Claude Haiku and the cheaper GPT tiers.