छवि संपादन और इनपेंटिंग
लिखित निर्देश से छवि के किसी हिस्से को बदलना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.
तकनीकी रूप से कैसे
The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.
प्रतिनिधि उत्पाद
5Firefly
2023रचनाकारों के लिए चित्र निर्माण और संपादन उपकरण
Stable Diffusion
2022टेक्स्ट-टू-इमेज वेट खुले किए और उन्हें उपभोक्ता GPU पर चलने लायक बनाया
FLUX
2024रेक्टिफाइड-फ्लो ट्रांसफॉर्मर से उच्च-रिज़ॉल्यूशन चित्र बनाता है
DALL·E 3
2023लंबे प्रॉम्प्ट को विस्तृत विवरण में बदलकर चित्र बनाता है
Seedream
2024मूल उच्च-रिज़ॉल्यूशन और मजबूत टेक्स्ट रेंडरिंग वाला टेक्स्ट-टू-इमेज मॉडल
संबंधित संस्थान
सामान्य उपयोग
- Removing clutter and swapping backgrounds in product photos
- Portrait retouching and restyling
- Swapping assets in ads and posters
- Restoring old photos and filling missing parts
इसका मूल्यांकन कैसे होता है
- Edit-direction consistency
- Whether the CLIP-space shift matches the instruction direction
- FID
- Distribution gap to real images, guarding against degrading realism
- Human rating
- Ratings for instruction completion and preservation
सीमाएँ और कठिनाइयाँ
- Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
- Multi-part requests — new outfit, new background, new expression — usually complete only some of them
- Large edits drift the identity, so the face stops resembling the original person
इसके पीछे की अवधारणाएँ
अव्यक्त-समष्टि विसरण और सशर्त नियंत्रण
विसरण पिक्सेल पर नहीं, बल्कि संपीड़ित अर्थ-समष्टि के भीतर चलाएँ
विसरण मॉडल
हज़ार छोटे शोर-हटाने के चरण सीखें, और शुद्ध शोर से चित्र बना सकेंगे
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है