Chuyển đến nội dung
Bản đồ AI

vLLM

Engine suy luận và phục vụ LLM thông lượng cao

UC Berkeley (LMSYS) Công cụ Mã nguồn mở

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

vLLM is an open-source inference and serving engine released in 2023 by a team at UC Berkeley. Its PagedAttention manages attention caches in pages to ease memory fragmentation, paired with continuous batching to raise throughput. It targets memory utilisation and concurrency efficiency in large-model serving.

Vì sao đáng ghi nhớ

PagedAttention solved paging of the KV cache, raising per-GPU throughput substantially and becoming a widely adopted base for open-source serving.

Thông số chính

Key technique
PagedAttention paged KV cache
Scheduling
Continuous batching
Interface
OpenAI-compatible serving API

Khái niệm liên quan

Sản phẩm cùng loại