WHAT IT IS
Mistral Large is the flagship model that France’s Mistral AI released in February 2024, offered under a commercial licence rather than open weights, complementing its open small models. Mistral AI was founded in 2023 in Paris by former DeepMind and Meta researchers. The model targets reasoning, multilingual work and function calling, with a context window in the 128K range. Mistral builds a community with a handful of open models, then serves enterprise needs with commercial ones such as Large.
Why it matters
It represents the dual-track strategy of open-sourcing small models to win a community and selling flagship models commercially, letting a non-US lab hold ground on both the open-weight and commercial sides.
Key specs
- Parameters
- 123B (Mistral Large 2)
- Context window
- 128K tokens
- Open weights
- No (commercial licence)
- Released
- 2024-02
Capabilities
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Tool Use & Function Calling
Let the model pick an API and fill its arguments
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Comparable products
Llama
2023The model family that made the open-weight route mainstream
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Command R
2024A commercial model built for retrieval augmentation and tool use
Phi
2023Small, efficient models built from curated data
Jamba
2024An open-weight model that mixes a state-space model with a Transformer
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Yi
2023A bilingual Chinese–English open-weight model with very-long-context versions
DBRX
2024Databricks’ open-weights MoE language model