MiniMax-M
An open-weight reasoning model with hybrid attention and a million-token context
WHAT IT IS
MiniMax-M1 is an open-weight reasoning model MiniMax released in June 2025. MiniMax was founded in Shanghai in 2021 by Yan Junjie and others. M1 uses a hybrid attention design that combines standard full-attention layers with linear-attention layers to cut the compute cost of long-sequence reasoning, and its stated context window reaches the million-token range. With 456B total parameters activating about 46B per token, it is a large-scale MoE reasoning model whose weights are released under an open licence. The company also runs generation product lines such as Hailuo.
Why it matters
It mixes linear and full attention and pushes context to a million tokens to cut the cost of long-sequence reasoning; among the larger open-weight reasoning models, it is a concrete landing of the open-weight camp on the reasoning track.
Key specs
- Parameters
- 456B (MoE, ~46B active)
- Context window
- 1M tokens
- Architecture
- Hybrid linear and full attention
- Open weights
- Yes
Capabilities
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Tool Use & Function Calling
Let the model pick an API and fill its arguments
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Comparable products
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Doubao
2023ByteDance’s general chat model and application
Step
2023A general model family aimed at multimodality and on-device use