मुख्य सामग्री पर जाएँ

टेक्स्ट-से-इमेज

एक वाक्य से छवि बनाना

इमेज जनरेशन और संपादनप्रारंभिक #18
इनपुटटेक्स्टइमेज

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.

तकनीकी रूप से कैसे

The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.

प्रतिनिधि उत्पाद

7

DALL·E 3

2023
OpenAI

लंबे प्रॉम्प्ट को विस्तृत विवरण में बदलकर चित्र बनाता है

मॉडल बंद स्रोत
टेक्स्टइमेज

Midjourney

2022
Midjourney

सौंदर्यपरक शैली के लिए जानी जाने वाली टेक्स्ट-टू-इमेज सेवा

ऐप बंद स्रोत
टेक्स्टइमेजइमेज

Stable Diffusion

2022
Stability AI

टेक्स्ट-टू-इमेज वेट खुले किए और उन्हें उपभोक्ता GPU पर चलने लायक बनाया

मॉडल खुले वेट
टेक्स्टइमेजइमेज

FLUX

2024
Black Forest Labs

रेक्टिफाइड-फ्लो ट्रांसफॉर्मर से उच्च-रिज़ॉल्यूशन चित्र बनाता है

मॉडल खुले वेट
टेक्स्टइमेजइमेज

Imagen

2022
Google DeepMind

कैस्केड विसरण और बड़े टेक्स्ट एनकोडर से यथार्थ चित्र बनाता है

मॉडल बंद स्रोत
टेक्स्टइमेज

Seedream

2024
ByteDance (Seed)

मूल उच्च-रिज़ॉल्यूशन और मजबूत टेक्स्ट रेंडरिंग वाला टेक्स्ट-टू-इमेज मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेज

Firefly

2023
Adobe

रचनाकारों के लिए चित्र निर्माण और संपादन उपकरण

ऐप बंद स्रोत
इमेजटेक्स्टइमेज

संबंधित संस्थान

सामान्य उपयोग

  • Concept design and storyboards
  • Marketing assets and illustration
  • Pre-visualisation for games and film
  • Personalised avatars and wallpapers

इसका मूल्यांकन कैसे होता है

FID
Distance to the real image distribution; lower is better
CLIP score
Semantic agreement between image and prompt
Human-preference Elo
Preference ranking from pairwise comparison

सीमाएँ और कठिनाइयाँ

  • Counting, precise spatial relations and ordering remain unreliable
  • Letters inside the image are frequently garbled, especially in long strings
  • Hands, limbs and object-contact points break down, often needing several resamples

इसके पीछे की अवधारणाएँ