テキストから画像生成
一文の説明から画像を一枚生成する
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.
技術的にどう実現するか
The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.
代表的な製品
7DALL·E 3
2023長い指示を詳細な説明に書き換えてから画像を生成する
Midjourney
2022美的スタイルで知られるテキストから画像生成サービス
Stable Diffusion
2022テキストから画像生成の重みを公開し、民生用GPUで動く規模に収めた
FLUX
2024整流フローTransformerで高解像度画像を生成する
Imagen
2022カスケード拡散と大規模テキストエンコーダで写実的な画像を生成する
Seedream
2024ネイティブ高解像度で文字描画に強いテキスト画像生成モデル
Firefly
2023クリエイター向けの画像生成・編集ツール
関連する組織
代表的な用途
- Concept design and storyboards
- Marketing assets and illustration
- Pre-visualisation for games and film
- Personalised avatars and wallpapers
どう評価するか
- FID
- Distance to the real image distribution; lower is better
- CLIP score
- Semantic agreement between image and prompt
- Human-preference Elo
- Preference ranking from pairwise comparison
限界と難しさ
- Counting, precise spatial relations and ordering remain unreliable
- Letters inside the image are frequently garbled, especially in long strings
- Hands, limbs and object-contact points break down, often needing several resamples