मुख्य सामग्री पर जाएँ

वाक् संश्लेषण

पाठ को स्वाभाविक आवाज़ में पढ़कर सुनाना

वाक् और संगीतप्रारंभिक #28
इनपुटटेक्स्टऑडियो

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes text and outputs audio that reads it aloud. It must handle both pronunciation — saying the right words — and prosody — where to pause and what to stress — while carrying a target speaker timbre. It runs opposite to speech recognition, and unlike voice cloning it does not need to imitate a specific person from a few samples but usually picks from preset voices.

तकनीकी रूप से कैसे

Modern systems are mostly two-stage: an acoustic model predicts an intermediate representation such as a mel spectrogram from text, then a vocoder restores the waveform; both can be trained end to end. Discrete audio tokens with autoregressive or diffusion decoding yield more natural prosody and support zero-shot timbre imitation. On the text side, phonemisation and text normalisation handle digits, abbreviations and polyphonic characters.

प्रतिनिधि उत्पाद

6

ElevenLabs

2022
ElevenLabs

स्वाभाविक और क्लोन करने योग्य आवाज़ वाली बहुभाषी टेक्स्ट-टू-स्पीच सेवा

API बंद स्रोत
टेक्स्टऑडियो

SparkTTS

2023
iFlytek

चीनी भाषा के उपयोग के लिए लक्षित टेक्स्ट-टू-स्पीच इंटरफ़ेस

API बंद स्रोत
टेक्स्टऑडियो

Synthesia

2019
Synthesia

टेक्स्ट लिखें, बोलते अवतार वाला वीडियो पाएँ

ऐप बंद स्रोत
टेक्स्टवीडियो

Gemini

2023
Google DeepMind

मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोवीडियोटेक्स्ट

Doubao

2023
ByteDance (Seed)

बाइटडांस का सामान्य संवाद मॉडल और ऐप

मॉडल बंद स्रोत
टेक्स्टइमेजटेक्स्ट

Siri

2011
Apple

वह शुरुआती वॉइस असिस्टेंट जिसने आम फ़ोन में आवाज़ से नियंत्रण लाया

ऐप बंद स्रोत
ऑडियोऑडियोटेक्स्ट

संबंधित संस्थान

सामान्य उपयोग

  • Audiobooks and news reading
  • Accessibility reading and assistive narration
  • Navigation prompts, podcasts and dubbing
  • Support and smart-device response audio

इसका मूल्यांकन कैसे होता है

MOS
Mean opinion score for naturalness
Intelligibility (WER)
Re-transcribing the output to measure how much is misheard
Speaker similarity
Closeness to the target timbre

सीमाएँ और कठिनाइयाँ

  • Pauses and stress in long sentences sound unnatural, making whole passages feel flat
  • Numbers, proper nouns and polyphonic characters are misread, and one normalisation slip derails the sentence
  • Control over emotion and style is coarse; only a few systems let you finely set pace and mood

इसके पीछे की अवधारणाएँ