本文へスキップ
AI図鑑

動画編集とリップシンク

映像や台詞を差し替え、口の動きを音声に合わせる

動画中級 #26
入力動画テキスト音声動画

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

この能力とは何か

Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.

技術的にどう実現するか

Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.

代表的な製品

5

関連する組織

代表的な用途

  • Multilingual dubbing and localisation
  • Corporate training and narrated courseware
  • Rapid re-versioning of ad creative
  • Fixing and erasing mistakes in post-production

どう評価するか

Lip-sync LSE-C / LSE-D
Consistency distance between audio and visual features; better when matched
Human rating
Subjective scores for lip naturalness and preservation
FVD
Overall distribution quality of the edited video

限界と難しさ

  • Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
  • Long edits accumulate flicker and identity drift, and the same person gradually deforms
  • Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth

背景にある概念