B200 (Blackwell)
A Blackwell-architecture GPU aimed at trillion-parameter models
WHAT IT IS
B200 is a Blackwell-architecture data-centre GPU NVIDIA announced in March 2024. Two dies are packaged together over a high-speed interconnect for about 208 billion transistors in total, with 192GB of HBM3e memory and support for lower-precision formats such as FP4 and FP6. It uses fifth-generation NVLink at 1.8 TB/s per GPU and targets training and inference of trillion-parameter models.
Why it matters
It brought memory to 192GB and supports low-precision formats such as FP4, succeeding H100 as the main hardware for a new generation of large-model training and inference and bringing the two-die package into data-centre GPUs.
Key specs
- Architecture
- Blackwell
- Memory
- 192GB HBM3e
- Transistors
- About 208 billion
- Interconnect
- NVLink 1.8 TB/s
- Released
- 2024-03
Related concepts
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Inference Optimization & Serving
Training happens once; inference happens a billion times a day — and serving is torn between fast first tokens and high throughput, which usually pull against each other