Texte vers image
Transformer une phrase en image
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.
Comment c'est fait
The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.
Produits représentatifs
7DALL·E 3
2023Réécrit une longue consigne en description détaillée, puis dessine l’image
Midjourney
2022Un service texte-image réputé pour son style esthétique
Stable Diffusion
2022A libéré les poids texte-image et les a rendus exécutables sur des GPU grand public
FLUX
2024Génère des images haute résolution avec un transformer à flux rectifié
Imagen
2022Génère des images photoréalistes par diffusion en cascade et un grand encodeur de texte
Seedream
2024Un modèle texte-image à haute résolution native et bon rendu du texte
Firefly
2023Un outil de génération et d’édition d’images pour les créateurs
Organisations concernées
Usages typiques
- Concept design and storyboards
- Marketing assets and illustration
- Pre-visualisation for games and film
- Personalised avatars and wallpapers
Comment on l'évalue
- FID
- Distance to the real image distribution; lower is better
- CLIP score
- Semantic agreement between image and prompt
- Human-preference Elo
- Preference ranking from pairwise comparison
Limites et points difficiles
- Counting, precise spatial relations and ordering remain unreliable
- Letters inside the image are frequently garbled, especially in long strings
- Hands, limbs and object-contact points break down, often needing several resamples
Concepts sous-jacents
Modèles de diffusion
Apprenez mille petits pas de débruitage et vous créerez une image à partir de bruit pur
Diffusion latente et contrôle conditionnel
Diffuser non pas sur les pixels, mais dans un espace sémantique comprimé
Vue d’ensemble des modèles génératifs
Les modèles discriminatifs répondent « qu’est-ce que c’est », les génératifs « à quoi cela devrait ressembler »