GPT-4o
A natively multimodal general model, with text, image and audio through one door
WHAT IT IS
GPT-4o is a general-purpose model OpenAI released in May 2024; the “o” stands for omni, meaning text, image and audio are handled inside one model. Vision and speech understanding share a single forward pass, and spoken conversation runs at near-real-time latency. Where earlier systems stitched several models into a pipeline, GPT-4o trains multimodality end to end and simplifies the interaction entry point. It is offered through both the API and ChatGPT, and serves as OpenAI’s main multimodal general model.
Why it matters
It moved multimodality from “several models stitched into a pipeline” to “one model end to end” and cut spoken-dialogue latency to near human-conversation levels; general models have since treated native multimodality as the default target.
Key specs
- Context window
- 128K tokens
- Modality
- Text, image, audio in; text, audio out
- Released
- 2024-05
- Open weights
- No
Capabilities
Conversation & Instruction Following
Understand intent across turns and act on it
Text Generation
Continue a passage, one word at a time
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Image Understanding & VQA
Look at an image and answer free-form questions
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Comparable products
Claude
2023A general chat model known for long context and safety alignment
Gemini
2023A natively multimodal general model built for very long context
Llama
2023The model family that made the open-weight route mainstream
Grok
2023A chat model tied to a social platform’s live data
Mistral Large
2024The flagship commercial model from a European open-weights lab
Command R
2024A commercial model built for retrieval augmentation and tool use
Kimi
2023A Chinese chat assistant known for long-context handling