Zum Inhalt springen
KI-Atlas

Vor- und Feintraining

Erst Sprache aus riesigem ungelabelten Text lernen, dann mit wenig Daten spezialisieren – das dateneffizienteste Paradigma der modernen KI

04 NLP und große SprachmodelleFortgeschritten5. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

Pretraining trains a general-purpose model on massive unlabelled text using a self-supervised objective; fine-tuning then adapts that model to a specific task or domain with little labelled data. Together they form a "general first, specialist later" two-stage paradigm that cuts the data needed per downstream task by orders of magnitude.

Intuition

Think of training a doctor: pretraining is general education, learning language, common sense and reasoning; fine-tuning is the specialisation, learned from a modest caseload. Without the general foundation, the specialist would need an unrealistic amount of data; with it, a few cases suffice — but drilling only on specialist cases can erode the general knowledge.

Abb. 1

A three-stage training stack: pretraining builds general knowledge, instruction tuning teaches compliance, alignment shapes preference — each stage builds on the one below

The higher the stage, the less data and the more specific the objective — yet every stage is capped by the quality of the pretraining beneath it.
Abb. 2

Scaling laws: with ample data, loss falls as a power law in parameters; once data runs out (a fixed 300B tokens, repeated), the curve flattens early or even turns up

  • Data-rich (Chinchilla-optimal)
  • Data-limited (300B tokens, repeated)

Funktionsweise

  1. 01

    Self-supervised objectives: masked vs autoregressive

    Both mainstream objectives manufacture labels from the text itself. Masked language modelling (BERT) hides random words for the model to recover and is naturally bidirectional; autoregressive modelling (GPT) predicts the next word and is naturally unidirectional. The first favours understanding, the second generation, and the choice fixes the ability profile.

  2. 02

    Scaling laws

    Experiments show test loss falls as a smooth power law in parameters, data and compute, holding across several orders of magnitude. This turns "should we scale further" from a gamble into an extrapolation — though the law only describes the trend; diminishing returns and data bottlenecks still need separate judgement.

  3. 03

    Efficient fine-tuning: full vs LoRA vs prompt tuning

    Full fine-tuning updates every parameter: the highest ceiling but heavy in memory and storage, needing a full copy of the weights per task. LoRA freezes the base model and attaches a pair of low-rank matrices beside selected weights, often cutting trainable parameters to a thousandth. Prompt tuning learns only a short trainable "soft prompt", the lightest change with the most limited expressiveness.

  4. 04

    Catastrophic forgetting and its remedies

    Continuing to train on a new task can rapidly degrade performance on old ones — catastrophic forgetting. Remedies include mixing in general data, updating only a small subset of parameters (as with LoRA), using a very small learning rate, and training added adapters while leaving the backbone untouched.

Kernformel

L(N, D) ≈ L∞ + a·N^(−α) + b·D^(−β)
The power-law form of scaling: loss falls smoothly with parameters N and data D, with exponents α and β. This explains why "just scale it" worked for a while.

Anwendungsfelder

  • Domain adaptation: turning a general model into a medical, legal or financial specialist with modest in-domain text
  • Instruction tuning: teaching the model to follow formats and requests via instruction–response pairs
  • Low-cost multi-tenant serving: one LoRA per customer on a shared base of weights
  • Continued pretraining: extending pretraining on a language or domain corpus to fill coverage gaps

Häufige Missverständnisse

  • When fine-tuning data is small and skewed the model overfits and visibly loses general ability. Catastrophic forgetting is not a theoretical worry but a routine outcome of small-data fine-tuning.
  • LoRA saves memory, but it is not equivalent to full fine-tuning. The further a task departs from the pretraining distribution and the more it needs to reshape internal representations, the more obvious LoRA’s shortfall.
  • Scaling laws give a trend, not a guarantee. They do not promise that bigger is always better: exhausting the data or repeating it can bend the curve flat or even upward early.

Schlüsselbegriffe

Self-supervised
Labels manufactured from the data itself, no manual annotation
Masked language modelling
Hide random words and recover them, a bidirectional objective
LoRA
Low-rank adapters training only a tiny number of new parameters
Catastrophic forgetting
Rapid loss of old abilities while learning a new task

Weiterführende Literatur