Speech Recognition (ASR)
Transcribe spoken audio into text
WHAT THIS CAPABILITY MEANS
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
How it is done
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Representative products
6Whisper
2022Transcribes speech in many languages and translates it into English
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Siri
2011The early voice assistant that brought spoken control to mainstream phones
Organizations involved
Typical uses
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
How it is evaluated
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Limits and hard parts
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Concepts behind it
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost