本文へスキップ
AI図鑑

Gemini

ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル

Google DeepMind モデル クローズド
入力テキスト画像音声動画テキスト

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

これは何か

Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.

なぜ覚えておく価値があるか

By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.

主な仕様

Context window
1M tokens (Gemini 1.5 Pro)
Modality
Text, image, audio, video in; text out
Released
2023-12
Open weights
No

対応する能力

関連する概念

同種の製品