음성 합성
글을 자연스러운 사람 목소리로 읽는다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.
기술적으로 구현하는 방법
Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.
대표 제품
6ElevenLabs
2022자연스럽고 복제 가능한 목소리의 다국어 음성 합성 서비스
SparkTTS
2023중국어 환경을 겨냥한 음성 합성 인터페이스
Synthesia
2019글자를 입력하면 아바타가 말하는 영상이 나온다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Doubao
2023바이트댄스의 범용 대화 모델과 앱
Siri
2011음성 비서를 주류 스마트폰에 퍼뜨린 초기 제품
관련 기관
대표적 용도
- Audiobooks and news reading
- Accessibility reading and assistive narration
- Navigation prompts, podcasts and dubbing
- Support and smart-device response audio
성능을 평가하는 방법
- MOS
- Mean opinion score for naturalness
- Intelligibility (WER)
- Re-transcribing the output to measure how much is misheard
- Speaker similarity
- Closeness to the target timbre
경계와 난점
- Pauses and stress in long sentences sound unnatural, making whole passages feel flat
- Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
- Control over emotion and style is coarse; only a few systems let you finely set pace and mood