Chuyển đến nội dung
Bản đồ AI

Gemini

Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài

Google DeepMind Mô hình Đóng
đầu vàoVăn bảnHình ảnhÂm thanhVideoVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.

Vì sao đáng ghi nhớ

By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.

Thông số chính

Context window
1M tokens (Gemini 1.5 Pro)
Modality
Text, image, audio, video in; text out
Released
2023-12
Open weights
No

Năng lực liên quan

Khái niệm liên quan

Sản phẩm cùng loại