DeepSeek-V3
An open-weight MoE with 671B parameters, activating 37B per token
WHAT IT IS
DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.
Why it matters
A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.
Key specs
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Training cost
- About US$5.58M (per the official technical report)
- Open weights
- Yes
Capabilities
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Code Generation
Write runnable code straight from a description
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Comparable products
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Llama
2023The model family that made the open-weight route mainstream
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Mistral Large
2024The flagship commercial model from a European open-weights lab
Baichuan
2023An open-weight general model aimed at Chinese-language use