画像編集とインペイント
一文の指示で画像の一部を書き換える
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.
技術的にどう実現するか
The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.
代表的な製品
5Firefly
2023クリエイター向けの画像生成・編集ツール
Stable Diffusion
2022テキストから画像生成の重みを公開し、民生用GPUで動く規模に収めた
FLUX
2024整流フローTransformerで高解像度画像を生成する
DALL·E 3
2023長い指示を詳細な説明に書き換えてから画像を生成する
Seedream
2024ネイティブ高解像度で文字描画に強いテキスト画像生成モデル
関連する組織
代表的な用途
- Removing clutter and swapping backgrounds in product photos
- Portrait retouching and restyling
- Swapping assets in ads and posters
- Restoring old photos and filling missing parts
どう評価するか
- Edit-direction consistency
- Whether the CLIP-space shift matches the instruction direction
- FID
- Distribution gap to real images, guarding against degrading realism
- Human rating
- Ratings for instruction completion and preservation
限界と難しさ
- Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
- Multi-part requests — new outfit, new background, new expression — usually complete only some of them
- Large edits drift the identity, so the face stops resembling the original person