Edição de vídeo e sincronia labial
Reeditar ou redublar para os lábios acompanharem o áudio
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
Como é feita tecnicamente
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
Produtos representativos
5Synthesia
2019Escreva um texto e receba um vídeo com um avatar que fala
SenseAvatar
2022Gera vídeo de humano digital com sincronia labial a partir de um retrato e uma faixa de voz
Kling
2024Um modelo de vídeo curto para texto-vídeo e imagem-vídeo
Hailuo
2024Um modelo de vídeo curto focado em seguir instruções e na linguagem de câmera
Veo
2024Gera clipes de vídeo em 1080p com planos coerentes
Organizações relacionadas
Usos típicos
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
Como avaliar se funciona bem
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
Limites e dificuldades
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth
Conceitos por trás
Geração multimodal
Um único modelo que aprende a falar, desenhar, mover-se e até modelar o mundo 3D
Modelos de difusão
Aprenda mil pequenos passos de remoção de ruído e criará uma imagem a partir de ruído puro
Difusão latente e controle condicional
Fazer difusão não sobre pixels, mas dentro de um espaço semântico comprimido