音声認識
話した音声を文字に書き起こす
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
技術的にどう実現するか
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
代表的な製品
6Whisper
2022多言語の音声を文字起こしし、英語へ翻訳する
GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
Siri
2011音声アシスタントを主流のスマホに広めた初期の存在
関連する組織
代表的な用途
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
どう評価するか
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
限界と難しさ
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search