音声合成
文字を自然な肉声で読み上げる
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
技術的にどう実現するか
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
代表的な製品
6ElevenLabs
2022自然でクローン可能な音声を提供する多言語音声合成サービス
SparkTTS
2023中国語シーンを想定した音声合成API
Synthesia
2019テキストを入力するとアバターが話す動画ができる
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Doubao
2023バイトダンスの汎用対話モデルとアプリ
Siri
2011音声アシスタントを主流のスマホに広めた初期の存在
関連する組織
代表的な用途
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
どう評価するか
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
限界と難しさ
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood