B200 (Blackwell)
GPU архитектуры Blackwell для моделей с триллионом параметров
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО
B200 is a Blackwell-architecture data-centre GPU NVIDIA announced in March 2024. Two dies are packaged together over a high-speed interconnect for about 208 billion transistors in total, with 192GB of HBM3e memory and support for lower-precision formats such as FP4 and FP6. It uses fifth-generation NVLink at 1.8 TB/s per GPU and targets training and inference of trillion-parameter models.
Почему стоит запомнить
It brought memory to 192GB and supports low-precision formats such as FP4, succeeding H100 as the main hardware for a new generation of large-model training and inference and bringing the two-die package into data-centre GPUs.
Ключевые характеристики
- Architecture
- Blackwell
- Memory
- 192GB HBM3e
- Transistors
- About 208 billion
- Interconnect
- NVLink 1.8 TB/s
- Released
- 2024-03
Связанные концепции
Инфраструктура обучения и вывода
Память определяет размер модели, а коммуникации — время обучения; чистая вычислительная мощность редко бывает узким местом
Оптимизация инференса и обслуживание
Обучение случается один раз, вывод — миллиарды раз в день; быстрый первый токен и высокая пропускная способность обычно тянут в разные стороны