過学習と正則化
モデルは訓練データを丸暗記したのに、新しいデータでは手も足も出ない
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
定義
Overfitting is when a model performs well on the training data but fails to generalise — it has memorised the noise and incidental detail of the training set instead of the underlying regularity. Regularisation is a family of techniques that deliberately constrain model complexity (penalising large weights, randomly dropping units, stopping early, augmenting data) in order to balance fit against generalisation.
直観的な理解
A student memorises every answer in the workbook, right down to a typo in the punctuation. On the same workbook they score perfectly; on a fresh exam they are lost. Memorising answers and learning the method are two different things. Regularisation is the teacher withholding the answer key and constantly reshuffling the problems — forcing the student to grasp the real pattern rather than recite.
Underfitting versus overfitting: two opposite ways of missing the sweet spot, with different symptoms and remedies
The classic shape of overfitting: training loss falls monotonically while validation loss turns upward after epoch 6; L2 regularisation flattens the rebound (click the legend to toggle curves)
- Training loss
- Validation loss
- Validation loss (with L2)
仕組み
- 01
Watch the two curves
Training loss keeps falling while validation loss falls, then rises. The point where the two curves begin to diverge is where overfitting starts.
- 02
Diagnose the current state
Both high means underfitting (model too weak); low training with high validation means overfitting (model too strong or too little data); both low and close is the sweet spot.
- 03
Apply constraints
Penalise large weights with L1/L2, randomly drop units with Dropout, stop before validation loss climbs, and widen the effective dataset with augmentation. The stronger the constraint, the lower the effective complexity.
- 04
Trade off on the validation set
Regularisation strength is a hyperparameter that can only be compared on the validation set. Too weak and overfitting persists; too strong and the model is crushed into underfitting. The balance must be measured, not derived.
重要公式
L_total = L_data + λ · Σ_w w²応用場面
- Deep learning: Dropout, weight decay and early stopping are near-standard
- Small-data modelling: fewer samples mean more overfitting, so regularisation and cross-validation matter most
- Tree models: limiting depth or minimum leaf size is a direct penalty on complexity
- Fine-tuning large models: low learning rates and light regularisation prevent catastrophic forgetting
よくある誤解
- Low training loss is not good news. It may simply be a symptom of overfitting; the only trustworthy signal is performance on the validation or test set.
- Regularisation is not free. Too much of it underfits the model and erases useful signal along with the noise; it must match the data volume and task complexity.
- Data volume matters more than regularisation. Adding real, diverse samples usually beats any regularisation trick; regularisation is merely compensation for insufficient data.
重要用語
- Generalisation
- Performance on data the model has not seen
- Weight decay (L2)
- Adding a squared-weight penalty to the loss to suppress large weights
- Dropout
- Randomly silencing units during training to prevent co-adaptation
- Early stopping
- Halting training before validation loss turns upward