Image vers vidéo
Mettre une image fixe en mouvement
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
Comment c'est fait
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
Produits représentatifs
8Stable Video Diffusion
2023Transforme une image fixe en courte vidéo avec un modèle de diffusion
Kling
2024Un modèle de vidéo courte pour texte-vidéo et image-vidéo
Hailuo
2024Un modèle de vidéo courte axé sur le suivi d’instructions et le langage caméra
Dream Machine
2024Génère de courtes vidéos animées à partir de texte ou d’une image
Veo
2024Génère des clips vidéo 1080p aux plans cohérents
Gen-3 Alpha
2024Un modèle texte-vidéo très contrôlable pour le cinéma et la publicité
Seedance
2024Un modèle de génération vidéo visant le récit multi-plans
Sora
2024Génère une vidéo cohérente d’environ une minute à partir d’une description
Organisations concernées
Usages typiques
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
Comment on l'évalue
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
Limites et points difficiles
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations
Concepts sous-jacents
Modèles de diffusion
Apprenez mille petits pas de débruitage et vous créerez une image à partir de bruit pur
Diffusion latente et contrôle conditionnel
Diffuser non pas sur les pixels, mais dans un espace sémantique comprimé
Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D