SenseAvatar
एक चेहरे की तस्वीर और ऑडियो से होंठ-सिंक डिजिटल-ह्यूमन वीडियो बनाता है
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
SenseAvatar is a digital-human video capability SenseTime introduced in August 2022 and offers as an interface. Given a portrait and a voice track, it generates a talking digital human whose lip movement matches the speech, with natural head motion. It is used for virtual presenters, announcements and customer service. The capability is a closed service.
यह क्यों महत्वपूर्ण है
It reduced digital-human generation to “one photo plus one voice track”, letting lip-synced avatar videos be produced at scale through an interface — an early example of digital humans moving toward large-scale use.
मुख्य विशिष्टताएँ
- Input
- A single portrait and a voice track
- Output
- Lip-synced talking digital-human video
- Availability
- API
- Open weights
- No
संबंधित क्षमताएँ
संबंधित अवधारणाएँ
ऑटोएन्कोडर और वीएई
सूचना को एक संकीर्णता से गुज़ारें, फिर उसे दोबारा उगने दें
जनरेटिव मॉडल: एक अवलोकन
विभेदक मॉडल बताते हैं "यह क्या है", जनरेटिव मॉडल बताते हैं "यह कैसा दिखना चाहिए"
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है