प्री-ट्रेनिंग और फाइन-ट्यूनिंग
पहले विशाल अलेबल पाठ से भाषा सीखना, फिर थोड़े डेटा से विशेषज्ञ बनना — आधुनिक एआई का सबसे डेटा-कुशल प्रतिमान
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
परिभाषा
Pretraining trains a general-purpose model on massive unlabelled text using a self-supervised objective; fine-tuning then adapts that model to a specific task or domain with little labelled data. Together they form a "general first, specialist later" two-stage paradigm that cuts the data needed per downstream task by orders of magnitude.
सहज समझ
Think of training a doctor: pretraining is general education, learning language, common sense and reasoning; fine-tuning is the specialisation, learned from a modest caseload. Without the general foundation, the specialist would need an unrealistic amount of data; with it, a few cases suffice — but drilling only on specialist cases can erode the general knowledge.
A three-stage training stack: pretraining builds general knowledge, instruction tuning teaches compliance, alignment shapes preference — each stage builds on the one below
Scaling laws: with ample data, loss falls as a power law in parameters; once data runs out (a fixed 300B tokens, repeated), the curve flattens early or even turns up
- Data-rich (Chinchilla-optimal)
- Data-limited (300B tokens, repeated)
कार्यप्रणाली
- 01
Self-supervised objectives: masked vs autoregressive
Both mainstream objectives manufacture labels from the text itself. Masked language modelling (BERT) hides random words for the model to recover and is naturally bidirectional; autoregressive modelling (GPT) predicts the next word and is naturally unidirectional. The first favours understanding, the second generation, and the choice fixes the ability profile.
- 02
Scaling laws
Experiments show test loss falls as a smooth power law in parameters, data and compute, holding across several orders of magnitude. This turns "should we scale further" from a gamble into an extrapolation — though the law only describes the trend; diminishing returns and data bottlenecks still need separate judgement.
- 03
Efficient fine-tuning: full vs LoRA vs prompt tuning
Full fine-tuning updates every parameter: the highest ceiling but heavy in memory and storage, needing a full copy of the weights per task. LoRA freezes the base model and attaches a pair of low-rank matrices beside selected weights, often cutting trainable parameters to a thousandth. Prompt tuning learns only a short trainable "soft prompt", the lightest change with the most limited expressiveness.
- 04
Catastrophic forgetting and its remedies
Continuing to train on a new task can rapidly degrade performance on old ones — catastrophic forgetting. Remedies include mixing in general data, updating only a small subset of parameters (as with LoRA), using a very small learning rate, and training added adapters while leaving the backbone untouched.
मुख्य सूत्र
L(N, D) ≈ L∞ + a·N^(−α) + b·D^(−β)उपयोग के क्षेत्र
- Domain adaptation: turning a general model into a medical, legal or financial specialist with modest in-domain text
- Instruction tuning: teaching the model to follow formats and requests via instruction–response pairs
- Low-cost multi-tenant serving: one LoRA per customer on a shared base of weights
- Continued pretraining: extending pretraining on a language or domain corpus to fill coverage gaps
सामान्य भ्रांतियाँ
- When fine-tuning data is small and skewed the model overfits and visibly loses general ability. Catastrophic forgetting is not a theoretical worry but a routine outcome of small-data fine-tuning.
- LoRA saves memory, but it is not equivalent to full fine-tuning. The further a task departs from the pretraining distribution and the more it needs to reshape internal representations, the more obvious LoRA’s shortfall.
- Scaling laws give a trend, not a guarantee. They do not promise that bigger is always better: exhausting the data or repeating it can bend the curve flat or even upward early.
मुख्य शब्द
- Self-supervised
- Labels manufactured from the data itself, no manual annotation
- Masked language modelling
- Hide random words and recover them, a bidirectional objective
- LoRA
- Low-rank adapters training only a tiny number of new parameters
- Catastrophic forgetting
- Rapid loss of old abilities while learning a new task