Редактирование видео и синхронизация губ
Перемонтировать или переозвучить видео под движение губ
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
Как это устроено
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
Примеры продуктов
5Synthesia
2019Введите текст — получите видео с говорящим цифровым аватаром
SenseAvatar
2022Создаёт видео цифрового человека с синхронизацией губ из портрета и дорожки голоса
Kling
2024Модель коротких видео как из текста, так и из изображения
Hailuo
2024Модель коротких видео с упором на следование инструкциям и язык камеры
Veo
2024Создаёт видеоклипы 1080p со связными планами
Связанные организации
Типичное применение
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
Как её оценивают
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
Границы и трудности
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth
Концепции в основе
Мультимодальная генерация
Одна модель учится говорить, рисовать, двигаться — и даже моделировать трёхмерный мир
Диффузионные модели
Научитесь тысяче мелких шагов удаления шума — и соберёте изображение из чистого шума
Диффузия в латентном пространстве и условное управление
Проводить диффузию не по пикселям, а внутри сжатого смыслового пространства