WHAT IT IS
vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.
Why it matters
PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.
Key specs
- Key technique
- PagedAttention paged KV cache
- Scheduling
- Continuous batching
- Interface
- OpenAI-compatible serving API
Related concepts
Inference Optimization & Serving
Training happens once; inference happens a billion times a day — and serving is torn between fast first tokens and high throughput, which usually pull against each other
Model Compression
Make a model smaller, faster and cheaper with almost no accuracy loss — but you can usually have only two of the three at once