Stable Diffusion
टेक्स्ट-टू-इमेज वेट खुले किए और उन्हें उपभोक्ता GPU पर चलने लायक बनाया
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
Stable Diffusion is an open-source text-to-image model that Stability AI released in August 2022. It runs the diffusion process in the latent space of a pretrained autoencoder, sharply cutting computation so the model can run on consumer GPUs. Text is turned into conditioning vectors by a CLIP text encoder and injected into the denoising network through cross-attention. The first 1.x version natively outputs 512×512, with a U-Net of roughly 860 million parameters and weights released under an open licence.
यह क्यों महत्वपूर्ण है
By open-sourcing text-to-image weights and making them runnable on consumer GPUs, it directly spawned the whole open-source image ecosystem of fine-tunes, plugins and local generation, marking the point where generative imagery reached a broad audience.
मुख्य विशिष्टताएँ
- Native resolution
- 512×512 (1.x)
- U-Net parameters
- About 860M
- Architecture
- Latent diffusion + CLIP text encoder
- Open weights
- Yes
- Released
- 2022-08
संबंधित क्षमताएँ
संबंधित अवधारणाएँ
अव्यक्त-समष्टि विसरण और सशर्त नियंत्रण
विसरण पिक्सेल पर नहीं, बल्कि संपीड़ित अर्थ-समष्टि के भीतर चलाएँ
विसरण मॉडल
हज़ार छोटे शोर-हटाने के चरण सीखें, और शुद्ध शोर से चित्र बना सकेंगे
ऑटोएन्कोडर और वीएई
सूचना को एक संकीर्णता से गुज़ारें, फिर उसे दोबारा उगने दें
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है
समान उत्पाद
Midjourney
2022सौंदर्यपरक शैली के लिए जानी जाने वाली टेक्स्ट-टू-इमेज सेवा
DALL·E 3
2023लंबे प्रॉम्प्ट को विस्तृत विवरण में बदलकर चित्र बनाता है
FLUX
2024रेक्टिफाइड-फ्लो ट्रांसफॉर्मर से उच्च-रिज़ॉल्यूशन चित्र बनाता है
Imagen
2022कैस्केड विसरण और बड़े टेक्स्ट एनकोडर से यथार्थ चित्र बनाता है
Firefly
2023रचनाकारों के लिए चित्र निर्माण और संपादन उपकरण
Seedream
2024मूल उच्च-रिज़ॉल्यूशन और मजबूत टेक्स्ट रेंडरिंग वाला टेक्स्ट-टू-इमेज मॉडल