이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
무엇인가
Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.
기억할 만한 이유
By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.
주요 사양
- Context window
- 1M tokens (Gemini 1.5 Pro)
- Modality
- Text, image, audio, video in; text out
- Released
- 2023-12
- Open weights
- No
소속 능력
관련 개념
동종 제품
GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Claude
2023긴 문맥과 안전 정렬로 알려진 범용 대화 모델
Llama
2023개방형 가중치 노선을 주류로 만든 범용 모델 계열
o3
2025답하기 전에 긴 추론을 거쳐, 추론 시 연산으로 정확도를 높인다
Grok
2023소셜 플랫폼 데이터와 결합된 대화 모델