Aller au contenu
Atlas de l'IA

SenseAvatar

Génère une vidéo d’humain numérique synchronisée sur les lèvres à partir d’un portrait et d’une piste vocale

SenseTime API Fermé
entréeImageAudioVidéo

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

CE QUE C'EST

SenseAvatar is a digital-human video capability SenseTime introduced in August 2022 and offers as an interface. Given a portrait and a voice track, it generates a talking digital human whose lip movement matches the speech, with natural head motion. It is used for virtual presenters, announcements and customer service. The capability is a closed service.

Pourquoi il compte

It reduced digital-human generation to “one photo plus one voice track”, letting lip-synced avatar videos be produced at scale through an interface — an early example of digital humans moving toward large-scale use.

Caractéristiques clés

Input
A single portrait and a voice track
Output
Lip-synced talking digital-human video
Availability
API
Open weights
No

Capacités

Concepts liés

Produits comparables