Montage vidéo et synchronisation labiale
Remonter ou redoubler la vidéo pour que les lèvres suivent
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
Comment c'est fait
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
Produits représentatifs
5Synthesia
2019Saisissez un texte, obtenez une vidéo avec un avatar qui parle
SenseAvatar
2022Génère une vidéo d’humain numérique synchronisée sur les lèvres à partir d’un portrait et d’une piste vocale
Kling
2024Un modèle de vidéo courte pour texte-vidéo et image-vidéo
Hailuo
2024Un modèle de vidéo courte axé sur le suivi d’instructions et le langage caméra
Veo
2024Génère des clips vidéo 1080p aux plans cohérents
Organisations concernées
Usages typiques
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
Comment on l'évalue
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
Limites et points difficiles
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth
Concepts sous-jacents
Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D
Modèles de diffusion
Apprenez mille petits pas de débruitage et vous créerez une image à partir de bruit pur
Diffusion latente et contrôle conditionnel
Diffuser non pas sur les pixels, mais dans un espace sémantique comprimé