Nhận dạng giọng nói
Chuyển giọng nói thành văn bản
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes an audio clip and outputs a written transcription, usually with timestamps and punctuation. It only turns sound into letters without interpreting it; it is the reverse direction of speech synthesis, and unlike video understanding it reads only the audio track. In multi-party recordings it must also separate speakers.
Làm ra sao về mặt kỹ thuật
Early systems were hybrids of an acoustic model and a language model, decoding frame by frame with hidden Markov models and a pronunciation lexicon. End-to-end models map acoustic features straight to characters or subwords, and Whisper-style training on large weakly labelled multilingual data markedly improved robustness to noise and accents. Streaming recognition must emit text before the audio ends, forcing a trade-off between latency and accuracy.
Sản phẩm tiêu biểu
6Whisper
2022Chuyển giọng nói nhiều ngôn ngữ thành văn bản và dịch sang tiếng Anh
GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
Hunyuan
2023Dòng mô hình phổ thông của Tencent, có bản trọng số mở
Siri
2011Trợ lý giọng nói sớm đưa điều khiển bằng lời nói vào điện thoại phổ thông
Tổ chức liên quan
Cách dùng tiêu biểu
- Automatic minutes for meetings and interviews
- Subtitle generation and video transcription
- Voice input and voice search
- Quality inspection and analytics of support calls
Đánh giá nó tốt hay không thế nào
- Word error rate
- Substitutions, deletions and insertions over the truth length
- Character error rate
- Preferred for languages without spaces, such as Chinese
- Real-time factor
- Processing time over audio duration, showing whether it keeps up live
Ranh giới và điểm khó
- Overlapping speech and heavy background noise make the error rate jump; crosstalk is nearly unresolvable
- Accents, dialects, jargon and rare proper nouns are the main sources of error
- Timestamps and punctuation are unstable on long audio, and bad segmentation hurts downstream reading and search
Các khái niệm đằng sau
Cơ chế chú ý (Attention)
Mọi vị trí đều có thể nhìn thẳng vào mọi vị trí khác và phân bổ chú ý theo mức liên quan
Kiến trúc Transformer
Thay cách truyền từng từ bằng một phòng họp nơi mọi từ cùng lên tiếng, để phụ thuộc xa chỉ còn cách một bước
Token hóa
Mô hình không đọc ký tự mà đọc token; cách tách từ âm thầm quyết định năng lực và chi phí