Whisper
Transcribes speech in many languages and translates it into English
WHAT IT IS
Whisper is a speech-recognition model OpenAI released in September 2022, with open weights and code. It was trained on roughly 680,000 hours of multilingual, multitask audio from the web and frames transcription, translation and language identification as a single sequence-to-sequence problem. The model is an encoder–decoder that takes 30-second audio segments as input. It shipped in several sizes from tiny to large, with the large version at about 1.55 billion parameters.
Why it matters
With open weights it made multilingual speech recognition broadly usable, becoming the default choice in many speech apps and transcription tools, and stands as a leading example of weak supervision with large-scale data in speech.
Key specs
- Parameters
- About 1.55B (large)
- Training data
- About 680,000 hours of audio
- Languages
- About 99
- Open weights
- Yes (MIT licence)
- Architecture
- Encoder–decoder transformer
Capabilities
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI