Video Editing & Lip Sync
Recut a video or re-voice it so lips match audio
WHAT THIS CAPABILITY MEANS
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
How it is done
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
Representative products
5Synthesia
2019Type text, get a talking-avatar video
SenseAvatar
2022Generates lip-synced digital-human video from a portrait and a voice track
Kling
2024A short-video model for both text-to-video and image-to-video
Hailuo
2024A short-video model focused on instruction following and camera language
Veo
2024Generates 1080p video clips with coherent shots
Organizations involved
Typical uses
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
How it is evaluated
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
Limits and hard parts
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth
Concepts behind it
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space