Síntesis de voz
Leer texto en voz alta con voz natural
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
Cómo se consigue técnicamente
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
Productos representativos
6ElevenLabs
2022Un servicio multilingüe de texto a voz con voces naturales y clonables
SparkTTS
2023Una interfaz de texto a voz orientada al uso en chino
Synthesia
2019Escribe texto y obtén un vídeo con un avatar que habla
Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Doubao
2023El modelo de chat general y la aplicación de ByteDance
Siri
2011El primer asistente de voz que llevó el control hablado a los teléfonos
Organizaciones relacionadas
Usos típicos
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
Cómo se evalúa
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
Límites y dificultades
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood
Conceptos detrás
Autoencoders y VAE
Comprime la información por un cuello de botella y deja que vuelva a crecer
Mecanismo de atención
Cada posición puede mirar directamente a todas las demás y repartir su atención según la relevancia
Arquitectura Transformer
Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto