이미지 편집과 인페인팅
한 문장 지시로 이미지의 일부를 바꾼다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.
기술적으로 구현하는 방법
The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.
대표 제품
5Firefly
2023크리에이터를 위한 이미지 생성·편집 도구
Stable Diffusion
2022텍스트-이미지 가중치를 공개하고 소비자용 GPU에서 돌아가게 만들었다
FLUX
2024정류 흐름 트랜스포머로 고해상도 이미지를 생성한다
DALL·E 3
2023긴 프롬프트를 상세한 설명으로 다시 쓴 뒤 이미지를 생성한다
Seedream
2024네이티브 고해상도와 뛰어난 문자 렌더링의 텍스트-이미지 모델
관련 기관
대표적 용도
- Removing clutter and swapping backgrounds in product photos
- Portrait retouching and restyling
- Swapping assets in ads and posters
- Restoring old photos and filling missing parts
성능을 평가하는 방법
- Edit-direction consistency
- Whether the CLIP-space shift matches the instruction direction
- FID
- Distribution gap to real images, guarding against degrading realism
- Human rating
- Ratings for instruction completion and preservation
경계와 난점
- Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
- Multi-part requests — new outfit, new background, new expression — usually complete only some of them
- Large edits drift the identity, so the face stops resembling the original person