Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
CUDA is the parallel computing platform and programming model introduced by NVIDIA in 2007 that lets developers use GPUs for general-purpose computing. It provides language extensions, a compiler, and acceleration libraries such as cuDNN and cuBLAS. It addresses efficiently mapping highly parallel computation such as neural-network training and inference onto GPUs.
Warum es wichtig ist
It has been the practical programming interface for AI compute for over a decade; the libraries and tools built around it form NVIDIA’s deepest ecosystem moat.
Wichtige Eckdaten
- First release
- 2007
- Positioning
- GPU general-purpose parallel computing platform and programming model
- Companion libraries
- cuDNN, cuBLAS and others
Verwandte Konzepte
Infrastruktur für Training und Inferenz
Der Speicher bestimmt, wie groß ein Modell sein darf, die Kommunikation, wie lange das Training dauert
Inferenz-Optimierung und Serving
Training passiert einmal, Inferenz milliardenfach am Tag – und erstes Token und Durchsatz stehen oft im Konflikt