이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
무엇인가
DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.
기억할 만한 이유
A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.
주요 사양
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Training cost
- About US$5.58M (per the official technical report)
- Open weights
- Yes
소속 능력
관련 개념
동종 제품
DeepSeek-R1
2025강화학습으로 추론 사슬을 훈련하고 MIT 라이선스로 가중치를 공개한 추론 모델
Llama
2023개방형 가중치 노선을 주류로 만든 범용 모델 계열
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
Mistral Large
2024유럽 오픈웨이트 연구소의 플래그십 상용 모델
Baichuan
2023중국어 환경을 겨냥한 오픈웨이트 범용 모델