मुख्य सामग्री पर जाएँ

वाक् पहचान

बोले गए ऑडियो को पाठ में बदलना

वाक् और संगीतप्रारंभिक #27
इनपुटऑडियोटेक्स्ट

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.

तकनीकी रूप से कैसे

Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.

प्रतिनिधि उत्पाद

6

Whisper

2022
OpenAI

कई भाषाओं की वाणी को टेक्स्ट में लिखता और अंग्रेज़ी में अनुवाद करता है

मॉडल खुले वेट
ऑडियोटेक्स्ट

GPT-4o

2024
OpenAI

मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोटेक्स्टऑडियो

Gemini

2023
Google DeepMind

मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोवीडियोटेक्स्ट

Qwen

2023
Alibaba (Qwen)

अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार

मॉडल खुले वेट
टेक्स्टइमेजटेक्स्ट

Hunyuan

2023
Tencent (Hunyuan)

ओपन-वेट संस्करणों सहित टेनसेंट का सामान्य मॉडल परिवार

मॉडल खुले वेट
टेक्स्टइमेजटेक्स्ट

Siri

2011
Apple

वह शुरुआती वॉइस असिस्टेंट जिसने आम फ़ोन में आवाज़ से नियंत्रण लाया

ऐप बंद स्रोत
ऑडियोऑडियोटेक्स्ट

संबंधित संस्थान

सामान्य उपयोग

  • Automatic minutes for meetings and interviews
  • Subtitle generation and video transcription
  • Voice input and voice search
  • Quality inspection and analytics of support calls

इसका मूल्यांकन कैसे होता है

Word error rate
Substitutions, deletions and insertions over the truth length
Character error rate
Preferred for languages without spaces, such as Chinese
Real-time factor
Processing time over audio duration, showing whether it keeps up live

सीमाएँ और कठिनाइयाँ

  • Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
  • Accents, dialects, jargon and rare proper nouns are the main sources of error
  • Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search

इसके पीछे की अवधारणाएँ