本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.
なぜ覚えておく価値があるか
By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.
主な仕様
- Context window
- 1M tokens (Gemini 1.5 Pro)
- Modality
- Text, image, audio, video in; text out
- Released
- 2023-12
- Open weights
- No
対応する能力
関連する概念
同種の製品
GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Llama
2023オープンウェイト路線を主流にした汎用モデル群
o3
2025答える前に長い推論を重ね、推論時の計算で正答率を上げる
Grok
2023ソーシャルプラットフォームのデータと結びついた対話モデル