Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.
Warum es wichtig ist
PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.
Wichtige Eckdaten
- Key technique
- PagedAttention paged KV cache
- Scheduling
- Continuous batching
- Interface
- OpenAI-compatible serving API
Verwandte Konzepte
Inferenz-Optimierung und Serving
Training passiert einmal, Inferenz milliardenfach am Tag – und erstes Token und Durchsatz stehen oft im Konflikt
Modellkompression
Ein Modell kleiner, schneller und günstiger machen, ohne Genauigkeit zu verlieren – aber meist nur zwei von drei zugleich