Saltar al contenido
Atlas de IA

Edición de vídeo y sincronía labial

Reeditar o cambiar la voz para que los labios coincidan

VídeoIntermedio #26
entradaVídeoTextoAudioVídeo

El texto completo se presenta en inglés; el título y el resumen están traducidos.

QUÉ SIGNIFICA ESTA CAPACIDAD

Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.

Cómo se consigue técnicamente

Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.

Productos representativos

5

Organizaciones relacionadas

Usos típicos

  • Multilingual dubbing and localisation
  • Corporate training and narrated courseware
  • Rapid re-versioning of ad creative
  • Fixing and erasing mistakes in post-production

Cómo se evalúa

Lip-sync LSE-C / LSE-D
Consistency distance between audio and visual features; better when matched
Human rating
Subjective scores for lip naturalness and preservation
FVD
Overall distribution quality of the edited video

Límites y dificultades

  • Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
  • Long edits accumulate flicker and identity drift, and the same person gradually deforms
  • Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth

Conceptos detrás