Изображение в изображение
Пересоздать похожее изображение из исходного
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an image and outputs a new one that resembles it in content but is rewritten in style or detail. The knob controlling similarity is usually noise strength: less noise stays close to the original, more noise approaches fresh creation. Unlike image editing it does not require a text instruction naming what to change; it restyles or reinterprets the whole image.
Как это устроено
The method encodes the input into latent space, adds noise for a chosen number of steps, and denoises back from that point, so the output keeps the structure while taking on a new style. Finer control comes from stacking conditioning branches: edges, depth, pose or a reference style each enter as extra conditions, which is especially common in tasks such as completion and line-art colourisation.
Примеры продуктов
6Stable Diffusion
2022Открыла веса генерации изображений и сделала их посильными для потребительских видеокарт
FLUX
2024Создаёт изображения высокого разрешения с помощью трансформера rectified flow
Firefly
2023Инструмент генерации и редактирования изображений для авторов
Diffusers
2022Единая реализация и планировщик для диффузионных моделей
Seedream
2024Модель текст-изображение с нативным высоким разрешением и хорошей отрисовкой текста
Midjourney
2022Сервис генерации изображений, известный своим эстетическим стилем
Связанные организации
Типичное применение
- Style transfer and photo stylisation
- Sketch and line-art colouring and completion
- Iterating from rough layouts to renders
- Series of variants on one subject
Как её оценивают
- FID
- Distance between results and the target distribution
- LPIPS perceptual distance
- Perceptual difference from the input image
- Structural consistency
- How well contours and subject placement follow the input
Границы и трудности
- Balancing structure retention against rewriting is hard; the same settings behave differently across images
- At higher strength the content drifts, and faces or text are the first details lost
- Wholesale redrawing wrecks layout, so posters and UI screenshots cannot be preserved as-is
Концепции в основе
Диффузионные модели
Научитесь тысяче мелких шагов удаления шума — и соберёте изображение из чистого шума
Диффузия в латентном пространстве и условное управление
Проводить диффузию не по пикселям, а внутри сжатого смыслового пространства
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел