B200 (Blackwell)
Eine GPU der Blackwell-Architektur für Modelle mit Billionen Parametern
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
B200 is a Blackwell-architecture data-centre GPU NVIDIA announced in March 2024. Two dies are packaged together over a high-speed interconnect for about 208 billion transistors in total, with 192GB of HBM3e memory and support for lower-precision formats such as FP4 and FP6. It uses fifth-generation NVLink at 1.8 TB/s per GPU and targets training and inference of trillion-parameter models.
Warum es wichtig ist
It brought memory to 192GB and supports low-precision formats such as FP4, succeeding H100 as the main hardware for a new generation of large-model training and inference and bringing the two-die package into data-centre GPUs.
Wichtige Eckdaten
- Architecture
- Blackwell
- Memory
- 192GB HBM3e
- Transistors
- About 208 billion
- Interconnect
- NVLink 1.8 TB/s
- Released
- 2024-03
Verwandte Konzepte
Infrastruktur für Training und Inferenz
Der Speicher bestimmt, wie groß ein Modell sein darf, die Kommunikation, wie lange das Training dauert
Inferenz-Optimierung und Serving
Training passiert einmal, Inferenz milliardenfach am Tag – und erstes Token und Durchsatz stehen oft im Konflikt