본문으로 건너뛰기
AI 도감

SenseAvatar

한 장의 인물 사진과 음성으로 입모양이 맞는 디지털 휴먼 영상을 생성한다

SenseTime API 클로즈드 소스
입력이미지오디오영상

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

SenseAvatar is a digital-human video capability SenseTime introduced in August 2022 and offers as an interface. Given a portrait and a voice track, it generates a talking digital human whose lip movement matches the speech, with natural head motion. It is used for virtual presenters, announcements and customer service. The capability is a closed service.

기억할 만한 이유

It reduced digital-human generation to “one photo plus one voice track”, letting lip-synced avatar videos be produced at scale through an interface — an early example of digital humans moving toward large-scale use.

주요 사양

Input
A single portrait and a voice track
Output
Lip-synced talking digital-human video
Availability
API
Open weights
No

소속 능력

관련 개념

동종 제품