Перейти к содержимому
Атлас ИИ

vLLM

Высокопроизводительный движок вывода и обслуживания LLM

UC Berkeley (LMSYS) Инструмент Открытый исходный код

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ЧТО ЭТО

vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.

Почему стоит запомнить

PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.

Ключевые характеристики

Key technique
PagedAttention paged KV cache
Scheduling
Continuous batching
Interface
OpenAI-compatible serving API

Связанные концепции

Аналоги