Text zu Bild
Aus einem Satz ein Bild machen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.
Wie sie technisch umgesetzt wird
The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.
Repräsentative Produkte
7DALL·E 3
2023Schreibt eine lange Eingabe in eine detaillierte Beschreibung um und zeichnet dann das Bild
Midjourney
2022Ein Text-zu-Bild-Dienst, bekannt für seinen ästhetischen Stil
Stable Diffusion
2022Veröffentlichte Text-zu-Bild-Gewichte offen und klein genug für Consumer-GPUs
FLUX
2024Erzeugt hochauflösende Bilder mit einem Rectified-Flow-Transformer
Imagen
2022Erzeugt fotorealistische Bilder mit kaskadierter Diffusion und großem Text-Encoder
Seedream
2024Ein Text-zu-Bild-Modell mit nativer Hochauflösung und guter Textdarstellung
Firefly
2023Ein Werkzeug zur Bilderzeugung und -bearbeitung für Gestaltende
Beteiligte Organisationen
Typische Verwendungen
- Concept design and storyboards
- Marketing assets and illustration
- Pre-visualisation for games and film
- Personalised avatars and wallpapers
Wie sie bewertet wird
- FID
- Distance to the real image distribution; lower is better
- CLIP score
- Semantic agreement between image and prompt
- Human-preference Elo
- Preference ranking from pairwise comparison
Grenzen und schwierige Punkte
- Counting, precise spatial relations and ordering remain unreliable
- Letters inside the image are frequently garbled, especially in long strings
- Hands, limbs and object-contact points break down, often needing several resamples
Konzepte dahinter
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Latente Diffusion und konditionale Steuerung
Diffusion nicht über Pixel, sondern in einem komprimierten semantischen Raum
Generative Modelle im Überblick
Diskriminative Modelle beantworten "Was ist das?", generative "Wie sollte das aussehen?"