B200 (Blackwell)
Uma GPU com arquitetura Blackwell voltada a modelos de trilhões de parâmetros
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE É
B200 is a Blackwell-architecture data-centre GPU NVIDIA announced in March 2024. Two dies are packaged together over a high-speed interconnect for about 208 billion transistors in total, with 192GB of HBM3e memory and support for lower-precision formats such as FP4 and FP6. It uses fifth-generation NVLink at 1.8 TB/s per GPU and targets training and inference of trillion-parameter models.
Por que vale a pena lembrar
It brought memory to 192GB and supports low-precision formats such as FP4, succeeding H100 as the main hardware for a new generation of large-model training and inference and bringing the two-die package into data-centre GPUs.
Especificações-chave
- Architecture
- Blackwell
- Memory
- 192GB HBM3e
- Transistors
- About 208 billion
- Interconnect
- NVLink 1.8 TB/s
- Released
- 2024-03
Conceitos relacionados
Infraestrutura de treinamento e inferência
A memória decide o tamanho do modelo treinável; a comunicação, quanto tempo leva
Otimização e serviço de inferência
O treino ocorre uma vez; a inferência, bilhões de vezes por dia — e o primeiro token e a vazão costumam se opor