यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
Mistral Large is the flagship model that France’s Mistral AI released in February 2024, offered under a commercial licence rather than open weights, complementing its open small models. Mistral AI was founded in 2023 in Paris by former DeepMind and Meta researchers. The model targets reasoning, multilingual work and function calling, with a context window in the 128K range. Mistral builds a community with a handful of open models, then serves enterprise needs with commercial ones such as Large.
यह क्यों महत्वपूर्ण है
It represents the dual-track strategy of open-sourcing small models to win a community and selling flagship models commercially, letting a non-US lab hold ground on both the open-weight and commercial sides.
मुख्य विशिष्टताएँ
- Parameters
- 123B (Mistral Large 2)
- Context window
- 128K tokens
- Open weights
- No (commercial licence)
- Released
- 2024-02
संबंधित क्षमताएँ
टेक्स्ट जनरेशन
पहले के पाठ को शब्द-दर-शब्द आगे बढ़ाता है
संवाद और निर्देश-पालन
बातचीत में मंशा समझकर उस पर अमल करना
तर्क और विचार-शृंखला
कठिन समस्या को मध्यवर्ती चरणों में बाँटकर हल करना
टूल और फंक्शन कॉल
मॉडल API चुने और उसके आर्गुमेंट भरे
संबंधित अवधारणाएँ
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
अटेंशन तंत्र
हर स्थान बाकी सभी स्थानों को सीधे देख सकता है और प्रासंगिकता के अनुसार ध्यान बाँट सकता है
प्री-ट्रेनिंग और फाइन-ट्यूनिंग
पहले विशाल अलेबल पाठ से भाषा सीखना, फिर थोड़े डेटा से विशेषज्ञ बनना — आधुनिक एआई का सबसे डेटा-कुशल प्रतिमान
समान उत्पाद
Llama
2023खुले भार के रास्ते को मुख्यधारा बनाने वाला मॉडल परिवार
GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Command R
2024पुनर्प्राप्ति-संवर्धन और टूल उपयोग के लिए बना वाणिज्यिक मॉडल
Phi
2023चुने हुए डेटा से बने छोटे और कुशल मॉडल
Jamba
2024स्टेट-स्पेस मॉडल और Transformer को मिलाने वाला ओपन-वेट मॉडल
DeepSeek-V3
2024671B पैरामीटर वाला ओपन-वेट MoE, प्रति टोकन केवल 37B सक्रिय
Yi
2023चीनी-अंग्रेज़ी द्विभाषी ओपन-वेट मॉडल, अति-लंबे संदर्भ संस्करणों के साथ
DBRX
2024Databricks का ओपन-वेट MoE भाषा मॉडल