음성 인식
말한 음성을 문자로 받아쓴다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
기술적으로 구현하는 방법
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
대표 제품
6Whisper
2022여러 언어의 음성을 받아쓰고 영어로 번역한다
GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
Siri
2011음성 비서를 주류 스마트폰에 퍼뜨린 초기 제품
관련 기관
대표적 용도
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
성능을 평가하는 방법
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
경계와 난점
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search