Jamba
An open-weight model that mixes a state-space model with a Transformer
WHAT IT IS
Jamba is an open-weight model AI21 Labs released in March 2024, interleaving Mamba state-space layers with Transformer attention layers inside one mixture-of-experts architecture. AI21 Labs was founded in Tel Aviv in 2017 by Amnon Shashua and others. This hybrid design holds long context while cutting memory and compute relative to a pure-attention model. Jamba ships as a 52B-parameter MoE that activates about 12B per token, with a context window in the 256K range.
Why it matters
It replaced most attention layers with state-space layers, testing whether a non-pure-Transformer architecture can save compute at long context — a concrete instance of hybrid architectures entering open-weight models.
Key specs
- Parameters
- 52B (MoE, ~12B active)
- Context window
- 256K tokens
- Architecture
- Mamba state-space layers interleaved with Transformer layers
- Open weights
- Yes
Capabilities
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Question Answering & RAG
Retrieve the evidence first, then answer from it
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Comparable products
Command R
2024A commercial model built for retrieval augmentation and tool use
Llama
2023The model family that made the open-weight route mainstream
Mistral Large
2024The flagship commercial model from a European open-weights lab