動画編集とリップシンク
映像や台詞を差し替え、口の動きを音声に合わせる
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
技術的にどう実現するか
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
代表的な製品
5Synthesia
2019テキストを入力するとアバターが話す動画ができる
SenseAvatar
2022一枚の人物画像と音声からリップシンクしたデジタルヒューマン動画を生成する
Kling
2024テキストからも画像からも動画を生成できるショート動画モデル
Hailuo
2024指示追従とカメラワークを重視したショート動画モデル
Veo
20241080pでショットの一貫した動画クリップを生成する
関連する組織
代表的な用途
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
どう評価するか
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
限界と難しさ
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth