Qwen
An open-weight family spanning many sizes, with multimodal versions
WHAT IT IS
Qwen is a language-model series Alibaba began open-sourcing in August 2023. It spans parameters from 0.5B to 72B and covers both text-only and multimodal versions, with the Qwen-VL line accepting image input. Alibaba releases base models alongside instruction-tuned versions and specialised variants for coding and mathematics. Generations such as Qwen2.5 support a context window in the 128K range, with weights downloadable under a permissive licence. Its complete ladder of sizes makes it a widely used base for secondary development in the open-source community.
Why it matters
With a full ladder from 0.5B to 72B plus both text and multimodal lines released openly, it lets researchers always find a size that fits their compute — one of the broadest catalogues in the open-weight ecosystem.
Key specs
- Parameters
- From 0.5B to 72B (Qwen2.5 family)
- Context window
- 128K tokens (Qwen2.5)
- Modality
- Text, image in (Qwen-VL); text out
- Open weights
- Yes
Capabilities
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Tool Use & Function Calling
Let the model pick an API and fill its arguments
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost
Comparable products
Llama
2023The model family that made the open-weight route mainstream
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Mistral Large
2024The flagship commercial model from a European open-weights lab
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Phi
2023Small, efficient models built from curated data
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Kimi
2023A Chinese chat assistant known for long-context handling
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Hunyuan
2023Tencent’s general model family, with open-weight versions
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context
Yi
2023A bilingual Chinese–English open-weight model with very-long-context versions
Baichuan
2023An open-weight general model aimed at Chinese-language use
DBRX
2024Databricks’ open-weights MoE language model