Edición de vídeo y sincronía labial
Reeditar o cambiar la voz para que los labios coincidan
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
Cómo se consigue técnicamente
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
Productos representativos
5Synthesia
2019Escribe texto y obtén un vídeo con un avatar que habla
SenseAvatar
2022Genera vídeo de humano digital con sincronía labial desde un retrato y una pista de voz
Kling
2024Un modelo de vídeo corto para texto a vídeo y imagen a vídeo
Hailuo
2024Un modelo de vídeo corto centrado en seguir instrucciones y el lenguaje de cámara
Veo
2024Genera clips de vídeo en 1080p con planos coherentes
Organizaciones relacionadas
Usos típicos
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
Cómo se evalúa
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
Límites y dificultades
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth
Conceptos detrás
Generación multimodal
Un mismo modelo que aprende a hablar, dibujar, moverse e incluso modelar el mundo 3D
Modelos de difusión
Aprende mil pequeños pasos de eliminación de ruido y podrás crear una imagen desde ruido puro
Difusión latente y control condicional
Hacer difusión no sobre píxeles, sino dentro de un espacio semántico comprimido