Распознавание речи
Перевести речь в текст
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Как это устроено
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Примеры продуктов
6Whisper
2022Расшифровывает речь на многих языках и переводит её на английский
GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
Siri
2011Ранний голосовой ассистент, принёсший управление голосом в массовые телефоны
Связанные организации
Типичное применение
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Как её оценивают
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Границы и трудности
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Концепции в основе
Механизм внимания
Каждая позиция может напрямую «видеть» все остальные и динамически распределять внимание по релевантности
Архитектура Transformer
Замена эстафеты по одному слову залом, где все говорят сразу, — и дальние зависимости оказываются в одном шаге
Токенизация
Модель читает не символы, а токены — способ разбиения тихо определяет и возможности, и стоимость