영상 편집과 립싱크
화면이나 대사를 바꾸고 입 모양을 음성에 맞춘다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a video plus a text instruction or a new audio track and outputs an edited version: lines can be replaced with matching lip movement, objects erased, or a person swapped for a synthetic avatar. Unlike image-to-video it does not generate fresh footage but makes local changes on an existing timeline while keeping it coherent.
기술적으로 구현하는 방법
Lip sync typically runs on two fronts: one branch drives face-region generation, redrawing the mouth from phonemes and audio features while preserving head pose and identity, and another aligns audio to video in time so it does not drift. Digital-human video models the person as a controllable 3D or neural representation driven by speech and expression parameters. Object removal borrows video completion, using neighbouring frames to fill the erased region consistently in time.
대표 제품
5Synthesia
2019글자를 입력하면 아바타가 말하는 영상이 나온다
SenseAvatar
2022한 장의 인물 사진과 음성으로 입모양이 맞는 디지털 휴먼 영상을 생성한다
Kling
2024텍스트와 이미지 모두로 영상을 만드는 숏폼 모델
Hailuo
2024지시 준수와 카메라 워크에 집중한 숏폼 모델
Veo
20241080p의 일관된 숏 영상을 생성한다
관련 기관
대표적 용도
- Multilingual dubbing and localisation
- Corporate training and narrated courseware
- Rapid re-versioning of ad creative
- Fixing and erasing mistakes in post-production
성능을 평가하는 방법
- Lip-sync LSE-C / LSE-D
- Consistency distance between audio and visual features; better when matched
- Human rating
- Subjective scores for lip naturalness and preservation
- FVD
- Overall distribution quality of the edited video
경계와 난점
- Fast speech, non-native pronunciation and plosive-heavy sentences desynchronise the lips most
- Long edits accumulate flicker and identity drift, and the same person gradually deforms
- Profile views, downward gaze and occlusion make face prediction unstable, exposing seams in the redrawn mouth