Chuyển đến nội dung
Bản đồ AI

Whisper

Chuyển giọng nói nhiều ngôn ngữ thành văn bản và dịch sang tiếng Anh

OpenAI Mô hình Trọng số mở
đầu vàoÂm thanhVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

Whisper is a speech-recognition model OpenAI released in September 2022, with open weights and code. It was trained on roughly 680,000 hours of multilingual, multitask audio from the web and frames transcription, translation and language identification as a single sequence-to-sequence problem. The model is an encoder–decoder that takes 30-second audio segments as input. It shipped in several sizes from tiny to large, with the large version at about 1.55 billion parameters.

Vì sao đáng ghi nhớ

With open weights it made multilingual speech recognition broadly usable, becoming the default choice in many speech apps and transcription tools, and stands as a leading example of weak supervision with large-scale data in speech.

Thông số chính

Parameters
About 1.55B (large)
Training data
About 680,000 hours of audio
Languages
About 99
Open weights
Yes (MIT licence)
Architecture
Encoder–decoder transformer

Năng lực liên quan

Khái niệm liên quan

Sản phẩm cùng loại