WHAT IT IS
Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.
Why it matters
By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.
Key specs
- Context window
- 1M tokens (Gemini 1.5 Pro)
- Modality
- Text, image, audio, video in; text out
- Released
- 2023-12
- Open weights
- No
Capabilities
Conversation & Instruction Following
Understand intent across turns and act on it
Text Generation
Continue a passage, one word at a time
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Image Understanding & VQA
Look at an image and answer free-form questions
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Comparable products
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Claude
2023A general chat model known for long context and safety alignment
Llama
2023The model family that made the open-weight route mainstream
o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
Grok
2023A chat model tied to a social platform’s live data