o3
It reasons at length before answering, trading inference-time compute for steadier accuracy
WHAT IT IS
o3 is a reasoning model OpenAI released in April 2025, part of the o series. Unlike models that answer directly, it performs an internal chain of reasoning before producing its output, then derives the answer from that reasoning. The series is trained with reinforcement learning so the model learns to spend more reasoning steps when they buy a more reliable answer. o3 accepts text and image input and serves as the main version for mathematics, coding and scientific reasoning tasks.
Why it matters
It turned test-time compute into a product: by reasoning at length before answering, o3 spends more compute at inference to buy steadier correctness, and made this the default shape of mainstream reasoning models.
Key specs
- Type
- Reasoning model
- Context window
- 200K tokens
- Modality
- Text, image in; text out
- Released
- 2025-04
Capabilities
Related concepts
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Comparable products
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Gemini
2023A natively multimodal general model built for very long context
Grok
2023A chat model tied to a social platform’s live data
Claude
2023A general chat model known for long context and safety alignment
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context