ضغط النماذج
تصغير النموذج وتسريعه وتخفيض تكلفته دون خسارة كبيرة في الدقة — لكن الجمع بين الثلاثة نادر
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Model compression is a family of techniques that reduce a model’s parameter count, memory footprint and compute while preserving accuracy. The main tools are quantisation (representing weights and activations at lower bit-widths), knowledge distillation (a small model learning from a large one’s outputs), pruning (removing redundant weights or structures) and low-rank factorisation (approximating a large matrix with several smaller ones).
الحدس المباشر
Packing a cluttered house into a compact flat: quantisation re-labels every item’s size from millimetres to centimetres — far fewer labels, same objects; pruning throws away what is never used; distillation has an apprentice copy the master’s finished work instead of repeating the master’s wandering path. Each frees space, and each can also quietly discard something small but essential.
A taxonomy of compression: quantisation changes bit-width, distillation swaps models, pruning deletes structure, factorisation lowers rank — and the four can be combined
Perplexity increase of different quantisation schemes on LLaMA-class models (illustrative magnitude, lower is better, FP16 as baseline): 4-bit is near-lossless on most tasks, while 2–3 bit degrades markedly
طريقة العمل
- 01
Quantise: fewer bits per number
Mapping an FP16 float to an INT8 or INT4 integer means first fixing a scale that covers the weights’ range. When weights are well behaved but activations have outliers, post-training quantisation (PTQ) suffices; otherwise quantisation-aware training (QAT) lets the model adapt to the low bit-width while it is still learning.
- 02
Distil: pour the teacher into the student
Train a small model (the student) on the soft probability outputs of a large one (the teacher). Unlike hard labels, soft labels carry information about how close one answer is to another, letting the student learn a smoother decision boundary at a fraction of the size.
- 03
Prune: find and remove redundancy
Score weights by magnitude or importance and remove the least impactful. Unstructured pruning deletes individual weights — high sparsity but not necessarily faster on commodity hardware; structured pruning removes whole neurons or attention heads — a smaller ratio but a real speed-up.
- 04
Factorise: two small matrices for one big one
If an m×n matrix has low rank, it factors into m×r and r×n pieces, cutting parameters from mn to r(m+n). SVD is the classic route; LoRA applies low-rank updates to fine-tuning — freezing the original weights and training only a small pair of matrices.
The cost–quality trade-off of compression schemes: x is relative inference cost, y is relative accuracy. The ideal region is top-left (low cost, high accuracy); hover for details
مجالات الاستخدام
- Edge and mobile deployment: fitting a model into a phone, vehicle or embedded device’s memory
- Cutting inference cost: INT4 weights let the same memory serve more concurrent requests
- Distilled small models: handling structured tasks with high-throughput, low-latency models
- Low-bit deployment of large models: running data-centre-scale models on consumer GPUs
مفاهيم خاطئة شائعة
- “4-bit is nearly lossless” is a common statistical outcome, not a guarantee. Some sensitive layers or long-tail tasks degrade visibly at low bit-width, so evaluation must be per-layer rather than uniform.
- Sparsity is not speed. A sparse matrix from unstructured pruning often sees no real gain on commodity GPUs because the access pattern is irregular, unless hardware or kernels support it explicitly.
- Distillation cannot exceed the teacher’s ceiling, and the student inherits the teacher’s biases and hallucinations. Compressing a large model into a small one preserves its flaws along with its abilities.
مصطلحات أساسية
- Quantisation
- Representing float weights and activations with low-bit integers
- Knowledge distillation
- Training a small model on a large model’s soft outputs
- Pruning
- Removing low-impact weights or whole structures
- Low-rank factorisation
- Approximating a large matrix by a product of two smaller ones