B200 (Blackwell)
Un GPU à architecture Blackwell visant les modèles à mille milliards de paramètres
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE C'EST
B200 is a Blackwell-architecture data-centre GPU NVIDIA announced in March 2024. Two dies are packaged together over a high-speed interconnect for about 208 billion transistors in total, with 192GB of HBM3e memory and support for lower-precision formats such as FP4 and FP6. It uses fifth-generation NVLink at 1.8 TB/s per GPU and targets training and inference of trillion-parameter models.
Pourquoi il compte
It brought memory to 192GB and supports low-precision formats such as FP4, succeeding H100 as the main hardware for a new generation of large-model training and inference and bringing the two-die package into data-centre GPUs.
Caractéristiques clés
- Architecture
- Blackwell
- Memory
- 192GB HBM3e
- Transistors
- About 208 billion
- Interconnect
- NVLink 1.8 TB/s
- Released
- 2024-03
Concepts liés
Infrastructure d’entraînement et d’inférence
La mémoire décide de la taille du modèle entraînable, la communication de la durée — le calcul brut est rarement le goulot
Optimisation et service d’inférence
L’entraînement a lieu une fois, l’inférence des milliards de fois par jour — et le premier token s’oppose souvent au débit