यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.
यह क्यों महत्वपूर्ण है
By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.
मुख्य विशिष्टताएँ
- Context window
- 1M tokens (Gemini 1.5 Pro)
- Modality
- Text, image, audio, video in; text out
- Released
- 2023-12
- Open weights
- No
संबंधित क्षमताएँ
संवाद और निर्देश-पालन
बातचीत में मंशा समझकर उस पर अमल करना
टेक्स्ट जनरेशन
पहले के पाठ को शब्द-दर-शब्द आगे बढ़ाता है
तर्क और विचार-शृंखला
कठिन समस्या को मध्यवर्ती चरणों में बाँटकर हल करना
छवि समझ और विज़ुअल प्रश्नोत्तर
छवि देखकर उस पर खुले प्रश्नों का उत्तर देना
संबंधित अवधारणाएँ
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
बहु-मॉडल जनरेशन
एक ही मॉडल बोलना, चित्र बनाना, हिलना, और यहाँ तक कि 3D संसार का नमूना बनाना सीखता है
अटेंशन तंत्र
हर स्थान बाकी सभी स्थानों को सीधे देख सकता है और प्रासंगिकता के अनुसार ध्यान बाँट सकता है
समान उत्पाद
GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Claude
2023लंबे संदर्भ और सुरक्षा-संरेखण के लिए जाना जाने वाला सामान्य संवाद मॉडल
Llama
2023खुले भार के रास्ते को मुख्यधारा बनाने वाला मॉडल परिवार
o3
2025उत्तर देने से पहले लंबी तर्क-श्रृंखला, अनुमान के समय गणना से अधिक सटीकता
Grok
2023सोशल प्लेटफ़ॉर्म के डेटा से जुड़ा संवाद मॉडल