मुख्य सामग्री पर जाएँ

छवि संपादन और इनपेंटिंग

लिखित निर्देश से छवि के किसी हिस्से को बदलना

इमेज जनरेशन और संपादनप्रारंभिक #20
इनपुटइमेजटेक्स्टइमेज

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.

तकनीकी रूप से कैसे

The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.

प्रतिनिधि उत्पाद

5

संबंधित संस्थान

सामान्य उपयोग

  • Removing clutter and swapping backgrounds in product photos
  • Portrait retouching and restyling
  • Swapping assets in ads and posters
  • Restoring old photos and filling missing parts

इसका मूल्यांकन कैसे होता है

Edit-direction consistency
Whether the CLIP-space shift matches the instruction direction
FID
Distribution gap to real images, guarding against degrading realism
Human rating
Ratings for instruction completion and preservation

सीमाएँ और कठिनाइयाँ

  • Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
  • Multi-part requests — new outfit, new background, new expression — usually complete only some of them
  • Large edits drift the identity, so the face stops resembling the original person

इसके पीछे की अवधारणाएँ