يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
Whisper is a speech-recognition model OpenAI released in September 2022, with open weights and code. It was trained on roughly 680,000 hours of multilingual, multitask audio from the web and frames transcription, translation and language identification as a single sequence-to-sequence problem. The model is an encoder–decoder that takes 30-second audio segments as input. It shipped in several sizes from tiny to large, with the large version at about 1.55 billion parameters.
لماذا يستحق التذكّر
With open weights it made multilingual speech recognition broadly usable, becoming the default choice in many speech apps and transcription tools, and stands as a leading example of weak supervision with large-scale data in speech.
المواصفات الأساسية
- Parameters
- About 1.55B (large)
- Training data
- About 680,000 hours of audio
- Languages
- About 99
- Open weights
- Yes (MIT licence)
- Architecture
- Encoder–decoder transformer
القدرات المرتبطة
المفاهيم ذات الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
التدريب المسبق والضبط الدقيق
تعلّم اللغة أولاً من نصوص ضخمة بلا وسوم ثم التخصص ببيانات قليلة — أكثر النماذج كفاءةً في البيانات