Chuyển đến nội dung
Bản đồ AI

DeepSeek-V3

MoE trọng số mở 671B tham số, chỉ kích hoạt 37B mỗi token

DeepSeek Mô hình Trọng số mở
đầu vàoVăn bảnVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.

Vì sao đáng ghi nhớ

A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.

Thông số chính

Parameters
671B (MoE, 37B active)
Context window
128K tokens
Training cost
About US$5.58M (per the official technical report)
Open weights
Yes

Năng lực liên quan

Khái niệm liên quan

Sản phẩm cùng loại