본문으로 건너뛰기
AI 도감

vLLM

고처리량 LLM 추론·서빙 엔진

UC Berkeley (LMSYS) 도구 오픈 소스

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.

기억할 만한 이유

PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.

주요 사양

Key technique
PagedAttention paged KV cache
Scheduling
Continuous batching
Interface
OpenAI-compatible serving API

관련 개념

동종 제품