التعرّف على الكلام
تحويل الكلام المسموع إلى نص
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
كيف تُنفَّذ تقنيًا
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
منتجات تمثيلية
6Whisper
2022ينسخ الكلام في لغات عدة ويكتبه نصًا ويترجمه إلى الإنجليزية
GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Gemini
2023نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
Hunyuan
2023عائلة نماذج تينسنت العامة، مع نسخ بأوزان مفتوحة
Siri
2011المساعد الصوتي المبكر الذي أدخل التحكم بالصوت إلى الهواتف الشائعة
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
كيف يُقاس مدى جودتها
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
الحدود والصعوبات
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
المفاهيم الكامنة وراءها
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
الترميز إلى رموز (Tokenization)
لا تقرأ النماذج الحروف بل الرموز، وطريقة التقسيم تحدّد القدرة والكلفة بصمت