Zum Inhalt springen
KI-Atlas

Modellkompression

Ein Modell kleiner, schneller und günstiger machen, ohne Genauigkeit zu verlieren – aber meist nur zwei von drei zugleich

08 KI-Engineering, Sicherheit und EthikFortgeschritten2. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

Model compression is a family of techniques that reduce a model’s parameter count, memory footprint and compute while preserving accuracy. The main tools are quantisation (representing weights and activations at lower bit-widths), knowledge distillation (a small model learning from a large one’s outputs), pruning (removing redundant weights or structures) and low-rank factorisation (approximating a large matrix with several smaller ones).

Intuition

Packing a cluttered house into a compact flat: quantisation re-labels every item’s size from millimetres to centimetres — far fewer labels, same objects; pruning throws away what is never used; distillation has an apprentice copy the master’s finished work instead of repeating the master’s wandering path. Each frees space, and each can also quietly discard something small but essential.

Abb. 1

A taxonomy of compression: quantisation changes bit-width, distillation swaps models, pruning deletes structure, factorisation lowers rank — and the four can be combined

Model compressionModel compressionQuantisationQuantisationPost-training quantisationPost-training quantisationQuantisation-aware trainingQuantisation-aware trainingINT8 / INT4 / AWQ / GPTQINT8 / INT4 / AWQ / GPTQKnowledge distillationKnowledge distillationLogit distillationLogit distillationFeature / hidden-state distillationFeature / hidden-state disti…PruningPruningUnstructured pruningUnstructured pruningStructured pruningStructured pruningLow-rank factorisationLow-rank factorisationSVDSVDLoRA-style adaptationLoRA-style adaptation
Abb. 2

Perplexity increase of different quantisation schemes on LLaMA-class models (illustrative magnitude, lower is better, FP16 as baseline): 4-bit is near-lossless on most tasks, while 2–3 bit degrades markedly

Funktionsweise

  1. 01

    Quantise: fewer bits per number

    Mapping an FP16 float to an INT8 or INT4 integer means first fixing a scale that covers the weights’ range. When weights are well behaved but activations have outliers, post-training quantisation (PTQ) suffices; otherwise quantisation-aware training (QAT) lets the model adapt to the low bit-width while it is still learning.

  2. 02

    Distil: pour the teacher into the student

    Train a small model (the student) on the soft probability outputs of a large one (the teacher). Unlike hard labels, soft labels carry information about how close one answer is to another, letting the student learn a smoother decision boundary at a fraction of the size.

  3. 03

    Prune: find and remove redundancy

    Score weights by magnitude or importance and remove the least impactful. Unstructured pruning deletes individual weights — high sparsity but not necessarily faster on commodity hardware; structured pruning removes whole neurons or attention heads — a smaller ratio but a real speed-up.

  4. 04

    Factorise: two small matrices for one big one

    If an m×n matrix has low rank, it factors into m×r and r×n pieces, cutting parameters from mn to r(m+n). SVD is the classic route; LoRA applies low-rank updates to fine-tuning — freezing the original weights and training only a small pair of matrices.

Abb. 4

The cost–quality trade-off of compression schemes: x is relative inference cost, y is relative accuracy. The ideal region is top-left (low cost, high accuracy); hover for details

Anwendungsfelder

  • Edge and mobile deployment: fitting a model into a phone, vehicle or embedded device’s memory
  • Cutting inference cost: INT4 weights let the same memory serve more concurrent requests
  • Distilled small models: handling structured tasks with high-throughput, low-latency models
  • Low-bit deployment of large models: running data-centre-scale models on consumer GPUs

Häufige Missverständnisse

  • “4-bit is nearly lossless” is a common statistical outcome, not a guarantee. Some sensitive layers or long-tail tasks degrade visibly at low bit-width, so evaluation must be per-layer rather than uniform.
  • Sparsity is not speed. A sparse matrix from unstructured pruning often sees no real gain on commodity GPUs because the access pattern is irregular, unless hardware or kernels support it explicitly.
  • Distillation cannot exceed the teacher’s ceiling, and the student inherits the teacher’s biases and hallucinations. Compressing a large model into a small one preserves its flaws along with its abilities.

Schlüsselbegriffe

Quantisation
Representing float weights and activations with low-bit integers
Knowledge distillation
Training a small model on a large model’s soft outputs
Pruning
Removing low-impact weights or whole structures
Low-rank factorisation
Approximating a large matrix by a product of two smaller ones

Weiterführende Literatur