본문으로 건너뛰기
AI 도감

사전학습과 미세조정

방대한 무표지 텍스트로 언어를 먼저 익히고 소량 데이터로 전문화한다 — 현대 AI에서 가장 데이터 효율적인 패러다임

04 자연어 처리와 대규모 언어 모델중급이 영역의 5번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Pretraining trains a general-purpose model on massive unlabelled text using a self-supervised objective; fine-tuning then adapts that model to a specific task or domain with little labelled data. Together they form a "general first, specialist later" two-stage paradigm that cuts the data needed per downstream task by orders of magnitude.

직관적 이해

Think of training a doctor: pretraining is general education, learning language, common sense and reasoning; fine-tuning is the specialisation, learned from a modest caseload. Without the general foundation, the specialist would need an unrealistic amount of data; with it, a few cases suffice — but drilling only on specialist cases can erode the general knowledge.

그림 1

A three-stage training stack: pretraining builds general knowledge, instruction tuning teaches compliance, alignment shapes preference — each stage builds on the one below

The higher the stage, the less data and the more specific the objective — yet every stage is capped by the quality of the pretraining beneath it.
그림 2

Scaling laws: with ample data, loss falls as a power law in parameters; once data runs out (a fixed 300B tokens, repeated), the curve flattens early or even turns up

  • Data-rich (Chinchilla-optimal)
  • Data-limited (300B tokens, repeated)

작동 원리

  1. 01

    Self-supervised objectives: masked vs autoregressive

    Both mainstream objectives manufacture labels from the text itself. Masked language modelling (BERT) hides random words for the model to recover and is naturally bidirectional; autoregressive modelling (GPT) predicts the next word and is naturally unidirectional. The first favours understanding, the second generation, and the choice fixes the ability profile.

  2. 02

    Scaling laws

    Experiments show test loss falls as a smooth power law in parameters, data and compute, holding across several orders of magnitude. This turns "should we scale further" from a gamble into an extrapolation — though the law only describes the trend; diminishing returns and data bottlenecks still need separate judgement.

  3. 03

    Efficient fine-tuning: full vs LoRA vs prompt tuning

    Full fine-tuning updates every parameter: the highest ceiling but heavy in memory and storage, needing a full copy of the weights per task. LoRA freezes the base model and attaches a pair of low-rank matrices beside selected weights, often cutting trainable parameters to a thousandth. Prompt tuning learns only a short trainable "soft prompt", the lightest change with the most limited expressiveness.

  4. 04

    Catastrophic forgetting and its remedies

    Continuing to train on a new task can rapidly degrade performance on old ones — catastrophic forgetting. Remedies include mixing in general data, updating only a small subset of parameters (as with LoRA), using a very small learning rate, and training added adapters while leaving the backbone untouched.

핵심 수식

L(N, D) ≈ L∞ + a·N^(−α) + b·D^(−β)
The power-law form of scaling: loss falls smoothly with parameters N and data D, with exponents α and β. This explains why "just scale it" worked for a while.

응용 분야

  • Domain adaptation: turning a general model into a medical, legal or financial specialist with modest in-domain text
  • Instruction tuning: teaching the model to follow formats and requests via instruction–response pairs
  • Low-cost multi-tenant serving: one LoRA per customer on a shared base of weights
  • Continued pretraining: extending pretraining on a language or domain corpus to fill coverage gaps

흔한 오해

  • When fine-tuning data is small and skewed the model overfits and visibly loses general ability. Catastrophic forgetting is not a theoretical worry but a routine outcome of small-data fine-tuning.
  • LoRA saves memory, but it is not equivalent to full fine-tuning. The further a task departs from the pretraining distribution and the more it needs to reshape internal representations, the more obvious LoRA’s shortfall.
  • Scaling laws give a trend, not a guarantee. They do not promise that bigger is always better: exhausting the data or repeating it can bend the curve flat or even upward early.

핵심 용어

Self-supervised
Labels manufactured from the data itself, no manual annotation
Masked language modelling
Hide random words and recover them, a bidirectional objective
LoRA
Low-rank adapters training only a tiny number of new parameters
Catastrophic forgetting
Rapid loss of old abilities while learning a new task

참고문헌