Перейти к содержимому
Атлас ИИ

Оптимизация инференса и обслуживание

Обучение случается один раз, вывод — миллиарды раз в день; быстрый первый токен и высокая пропускная способность обычно тянут в разные стороны

08 Инженерия, безопасность и этика ИИСреднийСтатья 3 в этой области

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ОПРЕДЕЛЕНИЕ

Inference optimization and serving is the engineering practice of letting a trained model answer user requests quickly, reliably and cheaply in production. It balances two conflicting metrics — per-request latency (especially the time to the first token) and system-wide throughput — using KV caching and paged attention, continuous batching, speculative decoding, and request routing with graceful degradation.

Интуиция

Picture a restaurant at peak hour. TTFT is how long a guest waits for the first dish; throughput is how many tables you serve all evening. Pre-serving a small appetiser calms one table but ties up the kitchen; cooking the same dish for many tables at once (batching) raises efficiency but makes every table wait longer. The real art of scheduling is keeping the kitchen full without driving guests away.

Рис. 1

The serving pipeline for one request: prefill and token-by-token decoding are two different workloads, with the scheduler re-forming batches between them

Рис. 2

How batch size affects latency and throughput at once (illustrative magnitude): throughput rises then saturates while TTFT grows roughly linearly — the two move in opposite directions on the same plot

  • Time to first token (ms)
  • Throughput (tokens/s)

Как это работает

  1. 01

    Prefill vs decode: two different workloads

    Processing the whole prompt (prefill) is a one-shot large matrix multiply, limited by compute; generating token by token (decode) does one token at a time, limited by memory bandwidth. Different bottlenecks call for different optimisations — which is why TTFT and per-token speed must be measured separately.

  2. 02

    KV cache and paging: stop recomputing the past

    Autoregressive generation needs every previous token’s Key and Value at each layer. Caching them shrinks each step from “recompute the whole history” to “compute only the new token”. But the KV cache grows linearly with context and concurrency, so PagedAttention borrows OS paging to manage it in blocks, cutting fragmentation and raising the concurrency ceiling.

  3. 03

    Continuous batching: keep the batch flowing

    Static batching holds resources until the longest sequence in the batch finishes. Continuous batching re-forms the batch every step, admitting a new request the moment another completes, so the GPU rarely idles — the main reason modern serving stacks outrun naive implementations by several times.

  4. 04

    Speculative decoding and fallback: speed versus safety

    Speculative decoding has a small model draft several tokens, which the large model verifies in one parallel pass, giving a real speed-up when the drafts hold. When traffic exceeds capacity, routing and fallback — sending requests to a smaller model, capping context, or queuing — decide whether the system degrades gracefully or collapses.

Рис. 4

Relative throughput of serving stacks on the same model and hardware (naive HuggingFace Transformers = 1×, illustrative magnitude): paging and continuous batching explain the order-of-magnitude gap

Области применения

  • Online chat: keeping TTFT within a few hundred milliseconds to preserve a fluid experience
  • High-concurrency API gateways: serving more users on the same hardware via continuous batching
  • Multi-model routing and fallback: choosing model sizes by request difficulty and budget
  • Capacity and cost planning: estimating hardware cost per thousand tokens at a given latency target

Частые заблуждения

  • Raising the batch size lifts throughput but raises every request’s latency. Throughput and latency are not the same objective, and a server must decide which one it is tuning for.
  • The KV cache is a major memory consumer. The longer the context and the higher the concurrency, the larger it grows; many “context does not fit” problems are really the cache exhausting memory, not the weights.
  • TTFT and subsequent-token speed have different bottlenecks. Techniques that raise decode throughput often fail to improve TTFT, and vice versa — they must be observed separately.

Ключевые термины

Time to first token (TTFT)
Time from sending a request to receiving the first token
KV cache
Caching past tokens’ keys and values to avoid recomputation
Continuous batching
Re-forming the batch every step to keep the GPU busy
Speculative decoding
A small model drafts and the large model verifies in parallel to speed up generation

Дополнительная литература