Reconocimiento de voz
Transcribir audio hablado a texto
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Cómo se consigue técnicamente
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Productos representativos
6Whisper
2022Transcribe voz en varios idiomas y la traduce al inglés
GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Hunyuan
2023La familia de modelos generales de Tencent, con versiones de pesos abiertos
Siri
2011El primer asistente de voz que llevó el control hablado a los teléfonos
Organizaciones relacionadas
Usos típicos
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Cómo se evalúa
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Límites y dificultades
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Conceptos detrás
Mecanismo de atención
Cada posición puede mirar directamente a todas las demás y repartir su atención según la relevancia
Arquitectura Transformer
Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto
Tokenización
Los modelos no leen caracteres, leen tokens; y cómo los dividas decide en silencio capacidad y coste