Saltar al contenido
Atlas de IA

Preentrenamiento y ajuste fino

Aprender el lenguaje primero con texto sin etiquetas y luego especializarse con pocos datos: el paradigma más eficiente en datos de la IA moderna

04 PLN y grandes modelos de lenguajeIntermedioEntrada 5 de este dominio

El texto completo se presenta en inglés; el título y el resumen están traducidos.

DEFINICIÓN

Pretraining trains a general-purpose model on massive unlabelled text using a self-supervised objective; fine-tuning then adapts that model to a specific task or domain with little labelled data. Together they form a "general first, specialist later" two-stage paradigm that cuts the data needed per downstream task by orders of magnitude.

Intuición

Think of training a doctor: pretraining is general education, learning language, common sense and reasoning; fine-tuning is the specialisation, learned from a modest caseload. Without the general foundation, the specialist would need an unrealistic amount of data; with it, a few cases suffice — but drilling only on specialist cases can erode the general knowledge.

Fig. 1

A three-stage training stack: pretraining builds general knowledge, instruction tuning teaches compliance, alignment shapes preference — each stage builds on the one below

The higher the stage, the less data and the more specific the objective — yet every stage is capped by the quality of the pretraining beneath it.
Fig. 2

Scaling laws: with ample data, loss falls as a power law in parameters; once data runs out (a fixed 300B tokens, repeated), the curve flattens early or even turns up

  • Data-rich (Chinchilla-optimal)
  • Data-limited (300B tokens, repeated)

Cómo funciona

  1. 01

    Self-supervised objectives: masked vs autoregressive

    Both mainstream objectives manufacture labels from the text itself. Masked language modelling (BERT) hides random words for the model to recover and is naturally bidirectional; autoregressive modelling (GPT) predicts the next word and is naturally unidirectional. The first favours understanding, the second generation, and the choice fixes the ability profile.

  2. 02

    Scaling laws

    Experiments show test loss falls as a smooth power law in parameters, data and compute, holding across several orders of magnitude. This turns "should we scale further" from a gamble into an extrapolation — though the law only describes the trend; diminishing returns and data bottlenecks still need separate judgement.

  3. 03

    Efficient fine-tuning: full vs LoRA vs prompt tuning

    Full fine-tuning updates every parameter: the highest ceiling but heavy in memory and storage, needing a full copy of the weights per task. LoRA freezes the base model and attaches a pair of low-rank matrices beside selected weights, often cutting trainable parameters to a thousandth. Prompt tuning learns only a short trainable "soft prompt", the lightest change with the most limited expressiveness.

  4. 04

    Catastrophic forgetting and its remedies

    Continuing to train on a new task can rapidly degrade performance on old ones — catastrophic forgetting. Remedies include mixing in general data, updating only a small subset of parameters (as with LoRA), using a very small learning rate, and training added adapters while leaving the backbone untouched.

Fórmula clave

L(N, D) ≈ L∞ + a·N^(−α) + b·D^(−β)
The power-law form of scaling: loss falls smoothly with parameters N and data D, with exponents α and β. This explains why "just scale it" worked for a while.

Dónde se usa

  • Domain adaptation: turning a general model into a medical, legal or financial specialist with modest in-domain text
  • Instruction tuning: teaching the model to follow formats and requests via instruction–response pairs
  • Low-cost multi-tenant serving: one LoRA per customer on a shared base of weights
  • Continued pretraining: extending pretraining on a language or domain corpus to fill coverage gaps

Errores comunes

  • When fine-tuning data is small and skewed the model overfits and visibly loses general ability. Catastrophic forgetting is not a theoretical worry but a routine outcome of small-data fine-tuning.
  • LoRA saves memory, but it is not equivalent to full fine-tuning. The further a task departs from the pretraining distribution and the more it needs to reshape internal representations, the more obvious LoRA’s shortfall.
  • Scaling laws give a trend, not a guarantee. They do not promise that bigger is always better: exhausting the data or repeating it can bend the curve flat or even upward early.

Términos clave

Self-supervised
Labels manufactured from the data itself, no manual annotation
Masked language modelling
Hide random words and recover them, a bidirectional objective
LoRA
Low-rank adapters training only a tiny number of new parameters
Catastrophic forgetting
Rapid loss of old abilities while learning a new task

Lecturas complementarias