画像から画像生成
元画像をもとに似た画像を生成し直す
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image and outputs a new one that resembles it in content but is rewritten in style or detail. The knob controlling similarity is usually noise strength: less noise stays close to the original, more noise approaches fresh creation. Unlike image editing it does not require a text instruction naming what to change; it restyles or reinterprets the whole image.
技術的にどう実現するか
The method encodes the input into latent space, adds noise for a chosen number of steps, and denoises back from that point, so the output keeps the structure while taking on a new style. Finer control comes from stacking conditioning branches: edges, depth, pose or a reference style each enter as extra conditions, which is especially common in tasks such as completion and line-art colourisation.
代表的な製品
6Stable Diffusion
2022テキストから画像生成の重みを公開し、民生用GPUで動く規模に収めた
FLUX
2024整流フローTransformerで高解像度画像を生成する
Firefly
2023クリエイター向けの画像生成・編集ツール
Diffusers
2022拡散モデルの統合実装とスケジューラ
Seedream
2024ネイティブ高解像度で文字描画に強いテキスト画像生成モデル
Midjourney
2022美的スタイルで知られるテキストから画像生成サービス
関連する組織
代表的な用途
- Style transfer and photo stylisation
- Sketch and line-art colouring and completion
- Iterating from rough layouts to renders
- Series of variants on one subject
どう評価するか
- FID
- Distance between results and the target distribution
- LPIPS perceptual distance
- Perceptual difference from the input image
- Structural consistency
- How well contours and subject placement follow the input
限界と難しさ
- Balancing structure retention against rewriting is hard; the same settings behave differently across images
- At higher strength the content drifts, and faces or text are the first details lost
- Wholesale redrawing wrecks layout, so posters and UI screenshots cannot be preserved as-is