Reconnaissance vocale
Transcrire la parole en texte
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Comment c'est fait
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Produits représentatifs
6Whisper
2022Transcrit la parole en plusieurs langues et la traduit en anglais
GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
Hunyuan
2023La famille de modèles généraux de Tencent, avec des versions à poids ouverts
Siri
2011L’assistant vocal qui a porté le contrôle à la voix sur les téléphones grand public
Organisations concernées
Usages typiques
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Comment on l'évalue
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Limites et points difficiles
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Concepts sous-jacents
Mécanisme d’attention
Chaque position peut regarder directement toutes les autres et répartir dynamiquement son attention selon la pertinence
Architecture Transformer
Remplacer le relais mot à mot par une salle où tous parlent en même temps, pour que les dépendances lointaines soient à un saut
Tokenisation
Les modèles ne lisent pas des caractères mais des tokens ; le découpage fixe silencieusement la capacité et le coût