本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.
なぜ覚えておく価値があるか
A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.
主な仕様
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Training cost
- About US$5.58M (per the official technical report)
- Open weights
- Yes
対応する能力
関連する概念
同種の製品
DeepSeek-R1
2025強化学習で推論の連鎖を訓練し、MIT ライセンスで重みを公開した推論モデル
Llama
2023オープンウェイト路線を主流にした汎用モデル群
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Mistral Large
2024欧州のオープンウェイト研究所による商用フラッグシップモデル
Baichuan
2023中国語シーンに向けたオープンウェイトの汎用モデル