Chuyển đến nội dung
Bản đồ AI

Tổng hợp giọng nói

Đọc văn bản thành giọng nói tự nhiên

Giọng nói & âm nhạcCơ bản #28
đầu vàoVăn bảnÂm thanh

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NĂNG LỰC NÀY NGHĨA LÀ GÌ

Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.

Làm ra sao về mặt kỹ thuật

Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.

Sản phẩm tiêu biểu

6

Tổ chức liên quan

Cách dùng tiêu biểu

  • Audiobooks and news reading
  • Accessibility reading and assistive narration
  • Navigation prompts, podcasts and dubbing
  • Support and smart-device response audio

Đánh giá nó tốt hay không thế nào

MOS
Mean opinion score for naturalness
Intelligibility (WER)
Re-transcribing the output to measure how much is misheard
Speaker similarity
Closeness to the target timbre

Ranh giới và điểm khó

  • Pauses and stress in long sentences sound unnatural, making whole passages feel flat
  • Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
  • Control over emotion and style is coarse; only a few systems let you finely set pace and mood

Các khái niệm đằng sau