Synthèse vocale
Lire un texte à voix haute naturellement
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
Comment c'est fait
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
Produits représentatifs
6ElevenLabs
2022Un service multilingue de synthèse vocale aux voix naturelles et clonables
SparkTTS
2023Une interface de synthèse vocale destinée aux usages en chinois
Synthesia
2019Saisissez un texte, obtenez une vidéo avec un avatar qui parle
Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
Doubao
2023Le modèle de conversation général et l’application de ByteDance
Siri
2011L’assistant vocal qui a porté le contrôle à la voix sur les téléphones grand public
Organisations concernées
Usages typiques
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
Comment on l'évalue
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
Limites et points difficiles
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood
Concepts sous-jacents
Autoencodeurs et VAE
Comprimer l’information dans un goulot, puis la laisser repousser
Mécanisme d’attention
Chaque position peut regarder directement toutes les autres et répartir dynamiquement son attention selon la pertinence
Architecture Transformer
Remplacer le relais mot à mot par une salle où tous parlent en même temps, pour que les dépendances lointaines soient à un saut