يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
SenseAvatar is a digital-human video capability SenseTime introduced in August 2022 and offers as an interface. Given a portrait and a voice track, it generates a talking digital human whose lip movement matches the speech, with natural head motion. It is used for virtual presenters, announcements and customer service. The capability is a closed service.
لماذا يستحق التذكّر
It reduced digital-human generation to “one photo plus one voice track”, letting lip-synced avatar videos be produced at scale through an interface — an early example of digital humans moving toward large-scale use.
المواصفات الأساسية
- Input
- A single portrait and a voice track
- Output
- Lip-synced talking digital-human video
- Availability
- API
- Open weights
- No
القدرات المرتبطة
المفاهيم ذات الصلة
المرمّزات الذاتية والمرمّزات الذاتية التباينية
اضغط المعلومات عبر عنق زجاجة ثم أعِد إنماءها
نظرة عامة على النماذج التوليدية
تجيب النماذج التمييزية عن «ما هذا»، بينما تجيب النماذج التوليدية عن «كيف ينبغي أن يبدو»
التوليد متعدد الوسائط
نموذج واحد يتعلّم الكلام والرسم والحركة، بل ونمذجة العالم ثلاثي الأبعاد