본문으로 건너뛰기
AI 도감

GPT-4o

네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다

OpenAI 모델 클로즈드 소스
입력텍스트이미지오디오텍스트오디오

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

GPT-4o is a general-purpose model OpenAI released in May 2024; the “o” stands for omni, meaning text, image and audio are handled inside one model. Vision and speech understanding share a single forward pass, and spoken conversation runs at near-real-time latency. Where earlier systems stitched several models into a pipeline, GPT-4o trains multimodality end to end and simplifies the interaction entry point. It is offered through both the API and ChatGPT, and serves as OpenAI’s main multimodal general model.

기억할 만한 이유

It moved multimodality from “several models stitched into a pipeline” to “one model end to end” and cut spoken-dialogue latency to near human-conversation levels; general models have since treated native multimodality as the default target.

주요 사양

Context window
128K tokens
Modality
Text, image, audio in; text, audio out
Released
2024-05
Open weights
No

소속 능력

관련 개념

동종 제품