Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE C'EST
vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.
Pourquoi il compte
PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.
Caractéristiques clés
- Key technique
- PagedAttention paged KV cache
- Scheduling
- Continuous batching
- Interface
- OpenAI-compatible serving API
Concepts liés
Optimisation et service d’inférence
L’entraînement a lieu une fois, l’inférence des milliards de fois par jour — et le premier token s’oppose souvent au débit
Compression de modèles
Rendre un modèle plus petit, plus rapide et moins cher sans perdre en précision — rarement les trois à la fois