本文へスキップ
AI図鑑

vLLM

高スループットのLLM推論・配信エンジン

UC Berkeley (LMSYS) ツール オープンソース

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

これは何か

vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.

なぜ覚えておく価値があるか

PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.

主な仕様

Key technique
PagedAttention paged KV cache
Scheduling
Continuous batching
Interface
OpenAI-compatible serving API

関連する概念

同種の製品