Image-to-Image
Regenerate a similar image from an input one
WHAT THIS CAPABILITY MEANS
Takes an image and outputs a new one that resembles it in content but is rewritten in style or detail. The knob controlling similarity is usually noise strength: less noise stays close to the original, more noise approaches fresh creation. Unlike image editing it does not require a text instruction naming what to change; it restyles or reinterprets the whole image.
How it is done
The method encodes the input into latent space, adds noise for a chosen number of steps, and denoises back from that point, so the output keeps the structure while taking on a new style. Finer control comes from stacking conditioning branches: edges, depth, pose or a reference style each enter as extra conditions, which is especially common in tasks such as completion and line-art colourisation.
Representative products
6Stable Diffusion
2022Released text-to-image weights openly and small enough to run on consumer GPUs
FLUX
2024Generates high-resolution images with a rectified-flow transformer
Firefly
2023An image generation and editing tool aimed at creators
Diffusers
2022A unified implementation and scheduler for diffusion models
Seedream
2024A text-to-image model with native high resolution and strong text rendering
Midjourney
2022A text-to-image service known for its aesthetic style
Organizations involved
Typical uses
- Style transfer and photo stylisation
- Sketch and line-art colouring and completion
- Iterating from rough layouts to renders
- Series of variants on one subject
How it is evaluated
- FID
- Distance between results and the target distribution
- LPIPS perceptual distance
- Perceptual difference from the input image
- Structural consistency
- How well contours and subject placement follow the input
Limits and hard parts
- Balancing structure retention against rewriting is hard; the same settings behave differently across images
- At higher strength the content drifts, and faces or text are the first details lost
- Wholesale redrawing wrecks layout, so posters and UI screenshots cannot be preserved as-is
Concepts behind it
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Image Representation
To a machine, a photo is nothing but stacked grids of numbers