본문으로 건너뛰기
AI 도감

Gemini

네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델

Google DeepMind 모델 클로즈드 소스
입력텍스트이미지오디오영상텍스트

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.

기억할 만한 이유

By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.

주요 사양

Context window
1M tokens (Gemini 1.5 Pro)
Modality
Text, image, audio, video in; text out
Released
2023-12
Open weights
No

소속 능력

관련 개념

동종 제품