Back to glossary

Term

Gemma 3

Gemma 3 is Googles family of open models released in March 2025. It comes in four sizes (1B, 4B, 12B, 27B), is multimodal from 4B upward, supports over 140 languages and offers a context window of up to 128,000 tokens.

Gemma 3 — explained in detail

Gemma 3 is a family of open-weight AI models that Google released on 12 March 2025. It is derived from the same research as Googles proprietary Gemini models but ships with freely downloadable weights, so it can be run locally or on your own infrastructure.

The family comes in four sizes: 1B, 4B, 12B and 27B parameters. The 1B model is text-only with a 32,000-token context; the three larger variants are multimodal, understanding images and short videos in addition to text, and offer a context window of up to 128,000 tokens. Gemma 3 supports over 140 languages as well as structured outputs and function calling for agentic workflows.

Technically, Gemma 3 uses an architecture that reduces the memory footprint of the KV cache by interleaving several local attention layers with a small window between global layers (roughly one global layer for every five local ones). This keeps the model resource-efficient despite its large context.

Example / practical relevance

Because the weights are openly available, Gemma 3 is often used where data protection, cost control or offline operation matter — for example in locally running assistants, in custom fine-tuning projects or on resource-constrained hardware. The small variants (1B, 4B) run on single GPUs or even on capable end devices.

Distinction

Unlike Googles Gemini models, which are only accessible via the API, Gemma 3 is an open-weight model for self-hosting. Against other open families (such as Llama or Qwen), Gemma 3 positions itself on multimodality, broad language coverage and strong performance at a comparatively compact size.

Discover more

Topic overview