ElevenLabs
A multilingual text-to-speech service with natural, cloneable voices
WHAT IT IS
ElevenLabs is a text-to-speech service the company launched in 2022 and offers through an API. It synthesises text into natural speech, supports multiple languages and voice cloning, and provides a low-latency streaming mode. Its multilingual model supports about 29 languages, and voice cloning lets users reproduce a specific timbre from a small number of samples. The service is provided in closed form.
Why it matters
It packaged near-human, timbre-controllable and cloneable speech synthesis into an easy interface, becoming one of the most widely called services for voice-cloning and dubbing applications.
Key specs
- Languages
- About 29 (multilingual model)
- Abilities
- Speech synthesis and voice cloning
- Low-latency mode
- Streaming synthesis
- Open weights
- No
Capabilities
Related concepts
Generative Models: An Overview
Discriminative models answer "what is this"; generative models answer "what should this look like"
Autoencoders & VAE
Squeeze information through a bottleneck, then let it grow back
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Comparable products
SparkTTS
2023A text-to-speech interface aimed at Chinese-language use
Whisper
2022Transcribes speech in many languages and translates it into English
Lyria
2023Generates instrumental and vocal music from text prompts
Suno
2023Generates complete songs with vocals from a single description