Generative Modelle für Bild- und Videoerstellung aus Text- oder Bild-Prompts — von Text-to-Image über Inpainting bis Text-to-Video. Anbieterübergreifende Kategorie, z. B. FLUX.1, Stable Diffusion, Sora und Veo.
FLUX.2 is the text-to-image model family from Black Forest Labs. It spans the proprietary Pro and Flex variants plus the open 32-billion-parameter Dev model and the compact Klein series, unifying image generation and image editing in a single model.
Google Imagen is the text-to-image model family from Google DeepMind. The current generation, Imagen 4, launched in 2025 in Fast, Generate and Ultra tiers, is available via the Gemini API and Vertex AI, and is known for strong typography and prompt adherence.
GPT Image 2 is OpenAIs native image model (April 2026) that reasons before drawing, renders text very reliably and produces high-resolution photorealistic images.
Kling 3.0 is Kuaishous video model, released on February 5, 2026. It generates photorealistic clips of up to 15 seconds with native audio across multiple languages and is regarded as the strongest price-to-performance video generator.
Midjourney v8 is the eighth generation of Midjourneys image generator (alpha March 2026) with roughly five times faster generation, native 2K resolution and improved text rendering.
Runway Gen-4.5 is Runways video model that turns text or an image into 5- and 10-second clips with strong prompt adherence and cinematic motion. It was built with NVIDIA and uses an Autoregressive-to-Diffusion technique.
Seedream 4.0 is ByteDances multimodal image model (September 2025) that processes text and multiple images as input, renders up to 4K and runs more than ten times faster than Seedream 3.0.
Veo 3.1 is Googles video model (DeepMind) that turns text or an image into 8-second clips with natively synchronized audio. Since the January 2026 update it delivers true 4K (3840x2160) and native vertical formats.
Redaktion·
We use cookies to improve your experience on this website. Privacy