यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
DALL·E 3 is OpenAI’s text-to-image model, released in October 2023. It folds prompt rewriting and image generation into one flow: the model first expands a short user prompt into a more detailed scene description, then generates the image from it, which improves adherence to long prompts and complex constraints. Training used automatic recaptioning to add fine-grained descriptions to images, narrowing the gap between images and their text labels. It is available through the API and ChatGPT.
यह क्यों महत्वपूर्ण है
It made “understand the prompt first, then draw” a single pipeline, markedly improving adherence to long prompts and constraints, and turned conversational image generation into a routine ChatGPT capability.
मुख्य विशिष्टताएँ
- Resolution
- 1024×1024, 1792×1024, 1024×1792
- Prompt handling
- Automatic recaptioning during training
- Open weights
- No
- Availability
- API and ChatGPT
संबंधित क्षमताएँ
संबंधित अवधारणाएँ
विसरण मॉडल
हज़ार छोटे शोर-हटाने के चरण सीखें, और शुद्ध शोर से चित्र बना सकेंगे
अव्यक्त-समष्टि विसरण और सशर्त नियंत्रण
विसरण पिक्सेल पर नहीं, बल्कि संपीड़ित अर्थ-समष्टि के भीतर चलाएँ
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है
प्रॉम्प्ट इंजीनियरिंग और अलाइनमेंट
मॉडल को उपयोगी, ईमानदार और हानिरहित बनाना उसे केवल बड़ा करने से कठिन है
समान उत्पाद
Midjourney
2022सौंदर्यपरक शैली के लिए जानी जाने वाली टेक्स्ट-टू-इमेज सेवा
Stable Diffusion
2022टेक्स्ट-टू-इमेज वेट खुले किए और उन्हें उपभोक्ता GPU पर चलने लायक बनाया
FLUX
2024रेक्टिफाइड-फ्लो ट्रांसफॉर्मर से उच्च-रिज़ॉल्यूशन चित्र बनाता है
Imagen
2022कैस्केड विसरण और बड़े टेक्स्ट एनकोडर से यथार्थ चित्र बनाता है
Firefly
2023रचनाकारों के लिए चित्र निर्माण और संपादन उपकरण
Seedream
2024मूल उच्च-रिज़ॉल्यूशन और मजबूत टेक्स्ट रेंडरिंग वाला टेक्स्ट-टू-इमेज मॉडल