Spracherkennung
Gesprochenes Audio in Text übertragen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Wie sie technisch umgesetzt wird
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Repräsentative Produkte
6Whisper
2022Transkribiert Sprache in vielen Sprachen und übersetzt sie ins Englische
GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Siri
2011Der frühe Sprachassistent, der Sprachsteuerung auf Massentelefone brachte
Beteiligte Organisationen
Typische Verwendungen
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Wie sie bewertet wird
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Grenzen und schwierige Punkte
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Konzepte dahinter
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt
Tokenisierung
Modelle lesen keine Zeichen, sondern Tokens – und die Zerlegung bestimmt still Fähigkeit wie Kosten