Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NÓ LÀ GÌ
H100 is a data-centre GPU NVIDIA announced in March 2022, based on the Hopper architecture. It carries 80GB of HBM3 memory, introduces fourth-generation tensor cores and a Transformer engine, and supports FP8 precision to raise throughput for large-model training and inference. Cards are linked by NVLink at 900 GB/s per GPU. H100 became one of the most widely used accelerators in large-model training clusters from 2022 to 2024.
Vì sao đáng ghi nhớ
It baked large-model-specific features such as FP8 and the Transformer engine into hardware, becoming the default accelerator in large-model training clusters after 2022 and a symbol of compute as the industry’s bottleneck.
Thông số chính
- Architecture
- Hopper
- Memory
- 80GB HBM3
- Process
- 4N (TSMC)
- Interconnect
- NVLink 900 GB/s
- Released
- 2022-03
Khái niệm liên quan
Hạ tầng huấn luyện và suy luận
Bộ nhớ quyết định mô hình lớn đến đâu, còn truyền thông quyết định mất bao lâu
Tối ưu suy luận và phục vụ
Huấn luyện chỉ một lần, suy luận diễn ra hàng tỷ lần mỗi ngày — và token đầu tiên lẫn thông lượng thường kéo ngược nhau