Image Editing & Inpainting
Change a spot in the image by a written instruction
WHAT THIS CAPABILITY MEANS
Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.
How it is done
The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.
Representative products
5Firefly
2023An image generation and editing tool aimed at creators
Stable Diffusion
2022Released text-to-image weights openly and small enough to run on consumer GPUs
FLUX
2024Generates high-resolution images with a rectified-flow transformer
DALL·E 3
2023Rewrites a long prompt into a detailed description, then draws the image
Seedream
2024A text-to-image model with native high resolution and strong text rendering
Organizations involved
Typical uses
- Removing clutter and swapping backgrounds in product photos
- Portrait retouching and restyling
- Swapping assets in ads and posters
- Restoring old photos and filling missing parts
How it is evaluated
- Edit-direction consistency
- Whether the CLIP-space shift matches the instruction direction
- FID
- Distribution gap to real images, guarding against degrading realism
- Human rating
- Ratings for instruction completion and preservation
Limits and hard parts
- Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
- Multi-part requests — new outfit, new background, new expression — usually complete only some of them
- Large edits drift the identity, so the face stops resembling the original person
Concepts behind it
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world