Bild zu Video
Ein Standbild in Bewegung bringen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
Wie sie technisch umgesetzt wird
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
Repräsentative Produkte
8Stable Video Diffusion
2023Macht aus einem Standbild mit einem Diffusionsmodell ein kurzes Video
Kling
2024Ein Kurzvideo-Modell für Text-zu-Video und Bild-zu-Video
Hailuo
2024Ein Kurzvideo-Modell mit Fokus auf Instruktionsbefolgung und Kamerasprache
Dream Machine
2024Erzeugt kurze Videos mit Bewegung aus Text oder einem Bild
Veo
2024Erzeugt 1080p-Videoclips mit kohärenten Einstellungen
Gen-3 Alpha
2024Ein gut steuerbares Text-zu-Video-Modell für Film und Werbung
Seedance
2024Ein Videogenerierungsmodell für mehrteiliges Erzählen
Sora
2024Erzeugt aus einer Beschreibung kohärentes Video von bis zu etwa einer Minute
Beteiligte Organisationen
Typische Verwendungen
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
Wie sie bewertet wird
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
Grenzen und schwierige Punkte
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations
Konzepte dahinter
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Latente Diffusion und konditionale Steuerung
Diffusion nicht über Pixel, sondern in einem komprimierten semantischen Raum
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt