Bildbearbeitung und Inpainting
Eine Stelle im Bild per Textanweisung ändern
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes the original image, a written instruction and usually a mask marking the region to change, and outputs the modified image. Unlike image-to-image it has an explicit, local edit intent and should leave unspecified areas untouched; unlike text-to-image it does not generate from scratch but operates surgically on an existing picture.
Wie sie technisch umgesetzt wird
The base method is inpainting: noise is added and removed only inside the mask while the rest reuses the original latents, with the instruction injected through cross-attention or an adapter. More careful schemes add a reference image and an identity-preservation branch so the edited person keeps the same face; another route hands instruction and image to a multimodal model that predicts the edited latents directly.
Repräsentative Produkte
5Firefly
2023Ein Werkzeug zur Bilderzeugung und -bearbeitung für Gestaltende
Stable Diffusion
2022Veröffentlichte Text-zu-Bild-Gewichte offen und klein genug für Consumer-GPUs
FLUX
2024Erzeugt hochauflösende Bilder mit einem Rectified-Flow-Transformer
DALL·E 3
2023Schreibt eine lange Eingabe in eine detaillierte Beschreibung um und zeichnet dann das Bild
Seedream
2024Ein Text-zu-Bild-Modell mit nativer Hochauflösung und guter Textdarstellung
Beteiligte Organisationen
Typische Verwendungen
- Removing clutter and swapping backgrounds in product photos
- Portrait retouching and restyling
- Swapping assets in ads and posters
- Restoring old photos and filling missing parts
Wie sie bewertet wird
- Edit-direction consistency
- Whether the CLIP-space shift matches the instruction direction
- FID
- Distribution gap to real images, guarding against degrading realism
- Human rating
- Ratings for instruction completion and preservation
Grenzen und schwierige Punkte
- Boundary, lighting and noise in the repainted area often mismatch the surroundings and need several passes
- Multi-part requests — new outfit, new background, new expression — usually complete only some of them
- Large edits drift the identity, so the face stops resembling the original person
Konzepte dahinter
Latente Diffusion und konditionale Steuerung
Diffusion nicht über Pixel, sondern in einem komprimierten semantischen Raum
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt