Sprachsynthese
Text mit natürlicher Stimme vorlesen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
Wie sie technisch umgesetzt wird
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
Repräsentative Produkte
6ElevenLabs
2022Ein mehrsprachiger Text-zu-Sprache-Dienst mit natürlichen, klonbaren Stimmen
SparkTTS
2023Eine Text-zu-Sprache-Schnittstelle für chinesischsprachige Anwendungen
Synthesia
2019Text eingeben, Video mit sprechendem Avatar erhalten
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Doubao
2023ByteDances allgemeines Dialogmodell und App
Siri
2011Der frühe Sprachassistent, der Sprachsteuerung auf Massentelefone brachte
Beteiligte Organisationen
Typische Verwendungen
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
Wie sie bewertet wird
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
Grenzen und schwierige Punkte
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood
Konzepte dahinter
Autoencoder und VAE
Information durch einen Engpass pressen und daraus wieder entstehen lassen
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt