Reconhecimento de fala
Transcrever áudio falado em texto
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Como é feita tecnicamente
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Produtos representativos
6Whisper
2022Transcreve fala em vários idiomas e traduz para o inglês
GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
Siri
2011O assistente de voz inicial que levou o controle por voz aos celulares populares
Organizações relacionadas
Usos típicos
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Como avaliar se funciona bem
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Limites e dificuldades
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Conceitos por trás
Mecanismo de atenção
Cada posição pode olhar diretamente para todas as outras e distribuir atenção conforme a relevância
Arquitetura Transformer
Substituir o revezamento palavra a palavra por uma sala onde todos falam ao mesmo tempo, deixando dependências distantes a um salto
Tokenização
Os modelos não leem caracteres, leem tokens; e como você divide define silenciosamente a capacidade e o custo