Text-to-Speech
Read text aloud as natural speech
WHAT THIS CAPABILITY MEANS
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
How it is done
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
Representative products
6ElevenLabs
2022A multilingual text-to-speech service with natural, cloneable voices
SparkTTS
2023A text-to-speech interface aimed at Chinese-language use
Synthesia
2019Type text, get a talking-avatar video
Gemini
2023A natively multimodal general model built for very long context
Doubao
2023ByteDance’s general chat model and application
Siri
2011The early voice assistant that brought spoken control to mainstream phones
Organizations involved
Typical uses
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
How it is evaluated
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
Limits and hard parts
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood
Concepts behind it
Autoencoders & VAE
Squeeze information through a bottleneck, then let it grow back
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away