SenseAvatar
Erzeugt lippensynchrones Digital-Human-Video aus einem Porträt und einer Tonspur
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
SenseAvatar is a digital-human video capability SenseTime introduced in August 2022 and offers as an interface. Given a portrait and a voice track, it generates a talking digital human whose lip movement matches the speech, with natural head motion. It is used for virtual presenters, announcements and customer service. The capability is a closed service.
Warum es wichtig ist
It reduced digital-human generation to “one photo plus one voice track”, letting lip-synced avatar videos be produced at scale through an interface — an early example of digital humans moving toward large-scale use.
Wichtige Eckdaten
- Input
- A single portrait and a voice track
- Output
- Lip-synced talking digital-human video
- Availability
- API
- Open weights
- No
Fähigkeiten
Verwandte Konzepte
Autoencoder und VAE
Information durch einen Engpass pressen und daraus wieder entstehen lassen
Generative Modelle im Überblick
Diskriminative Modelle beantworten "Was ist das?", generative "Wie sollte das aussehen?"
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt