Term
Google Imagen
Google Imagen is the text-to-image model family from Google DeepMind. The current generation, Imagen 4, launched in 2025 in Fast, Generate and Ultra tiers, is available via the Gemini API and Vertex AI, and is known for strong typography and prompt adherence.
Google Imagen — explained in more detail
Google Imagen is the text-to-image model family from Google DeepMind. The current generation, Imagen 4, was unveiled at Google I/O in 2025 and turns a text prompt into photorealistic images with high detail fidelity. Its notable strengths are precise text rendering (typography) within images, strong prompt adherence across many styles, multilingual prompt support and output up to 2K resolution.
Imagen 4 comes in three tiers: Fast (speed-optimized), Generate (standard) and Ultra (maximum prompt alignment). It is a closed, hosted model with no open weights; access runs through the Gemini API, Google AI Studio and Vertex AI and is billed per generated image. Google is increasingly routing generative image work to its Gemini-native models (the Nano Banana line).
Example / In practice
A marketing team uses Imagen 4 via Vertex AI to produce campaign visuals with legible embedded text — such as slogans or price tags — an area where image models have traditionally struggled. The Fast tier serves rapid iteration, while the Ultra tier is used for final assets.
Distinction from similar terms
Imagen is an image model and falls under the image & video use case; for video Google relies on Veo. Unlike open families such as FLUX or Stable Diffusion, Imagen is a closed API model and cannot be self-hosted.