Chuyển đến nội dung
Bản đồ AI

DeepSeek-R1

Mô hình suy luận huấn luyện bằng RL trên chuỗi suy nghĩ, trọng số mở theo giấy phép MIT

DeepSeek Mô hình Trọng số mở
đầu vàoVăn bảnVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.

Vì sao đáng ghi nhớ

It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.

Thông số chính

Parameters
671B (MoE, 37B active)
Context window
128K tokens
Open weights
Yes (MIT licence)
Type
Reasoning model

Năng lực liên quan

Khái niệm liên quan

Sản phẩm cùng loại