본문으로 건너뛰기
AI 도감

Whisper

여러 언어의 음성을 받아쓰고 영어로 번역한다

OpenAI 모델 공개 가중치
입력오디오텍스트

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

Whisper is a speech-recognition model OpenAI released in September 2022, with open weights and code. It was trained on roughly 680,000 hours of multilingual, multitask audio from the web and frames transcription, translation and language identification as a single sequence-to-sequence problem. The model is an encoder–decoder that takes 30-second audio segments as input. It shipped in several sizes from tiny to large, with the large version at about 1.55 billion parameters.

기억할 만한 이유

With open weights it made multilingual speech recognition broadly usable, becoming the default choice in many speech apps and transcription tools, and stands as a leading example of weak supervision with large-scale data in speech.

주요 사양

Parameters
About 1.55B (large)
Training data
About 680,000 hours of audio
Languages
About 99
Open weights
Yes (MIT licence)
Architecture
Encoder–decoder transformer

소속 능력

관련 개념

동종 제품