मुख्य सामग्री पर जाएँ

इमेज-से-इमेज

एक छवि से मिलती-जुलती छवि दोबारा बनाना

इमेज जनरेशन और संपादनप्रारंभिक #19
इनपुटइमेजइमेज

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes an image and outputs a new one that resembles it in content but is rewritten in style or detail. The knob controlling similarity is usually noise strength: less noise stays close to the original, more noise approaches fresh creation. Unlike image editing it does not require a text instruction naming what to change; it restyles or reinterprets the whole image.

तकनीकी रूप से कैसे

The method encodes the input into latent space, adds noise for a chosen number of steps, and denoises back from that point, so the output keeps the structure while taking on a new style. Finer control comes from stacking conditioning branches: edges, depth, pose or a reference style each enter as extra conditions, which is especially common in tasks such as completion and line-art colourisation.

प्रतिनिधि उत्पाद

6

Stable Diffusion

2022
Stability AI

टेक्स्ट-टू-इमेज वेट खुले किए और उन्हें उपभोक्ता GPU पर चलने लायक बनाया

मॉडल खुले वेट
टेक्स्टइमेजइमेज

FLUX

2024
Black Forest Labs

रेक्टिफाइड-फ्लो ट्रांसफॉर्मर से उच्च-रिज़ॉल्यूशन चित्र बनाता है

मॉडल खुले वेट
टेक्स्टइमेजइमेज

Firefly

2023
Adobe

रचनाकारों के लिए चित्र निर्माण और संपादन उपकरण

ऐप बंद स्रोत
इमेजटेक्स्टइमेज

Diffusers

2022
Hugging Face

डिफ्यूज़न मॉडल हेतु एकीकृत कार्यान्वयन व शेड्यूलर

टूल ओपन सोर्स

Seedream

2024
ByteDance (Seed)

मूल उच्च-रिज़ॉल्यूशन और मजबूत टेक्स्ट रेंडरिंग वाला टेक्स्ट-टू-इमेज मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेज

Midjourney

2022
Midjourney

सौंदर्यपरक शैली के लिए जानी जाने वाली टेक्स्ट-टू-इमेज सेवा

ऐप बंद स्रोत
टेक्स्टइमेजइमेज

संबंधित संस्थान

सामान्य उपयोग

  • Style transfer and photo stylisation
  • Sketch and line-art colouring and completion
  • Iterating from rough layouts to renders
  • Series of variants on one subject

इसका मूल्यांकन कैसे होता है

FID
Distance between results and the target distribution
LPIPS perceptual distance
Perceptual difference from the input image
Structural consistency
How well contours and subject placement follow the input

सीमाएँ और कठिनाइयाँ

  • Balancing structure retention against rewriting is hard; the same settings behave differently across images
  • At higher strength the content drifts, and faces or text are the first details lost
  • Wholesale redrawing wrecks layout, so posters and UI screenshots cannot be preserved as-is

इसके पीछे की अवधारणाएँ