本文へスキップ
AI図鑑

過学習と正則化

モデルは訓練データを丸暗記したのに、新しいデータでは手も足も出ない

02 機械学習中級この領域の第 4 項目

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

Overfitting is when a model performs well on the training data but fails to generalise — it has memorised the noise and incidental detail of the training set instead of the underlying regularity. Regularisation is a family of techniques that deliberately constrain model complexity (penalising large weights, randomly dropping units, stopping early, augmenting data) in order to balance fit against generalisation.

直観的な理解

A student memorises every answer in the workbook, right down to a typo in the punctuation. On the same workbook they score perfectly; on a fresh exam they are lost. Memorising answers and learning the method are two different things. Regularisation is the teacher withholding the answer key and constantly reshuffling the problems — forcing the student to grasp the real pattern rather than recite.

図 1

Underfitting versus overfitting: two opposite ways of missing the sweet spot, with different symptoms and remedies

図 2

The classic shape of overfitting: training loss falls monotonically while validation loss turns upward after epoch 6; L2 regularisation flattens the rebound (click the legend to toggle curves)

  • Training loss
  • Validation loss
  • Validation loss (with L2)

仕組み

  1. 01

    Watch the two curves

    Training loss keeps falling while validation loss falls, then rises. The point where the two curves begin to diverge is where overfitting starts.

  2. 02

    Diagnose the current state

    Both high means underfitting (model too weak); low training with high validation means overfitting (model too strong or too little data); both low and close is the sweet spot.

  3. 03

    Apply constraints

    Penalise large weights with L1/L2, randomly drop units with Dropout, stop before validation loss climbs, and widen the effective dataset with augmentation. The stronger the constraint, the lower the effective complexity.

  4. 04

    Trade off on the validation set

    Regularisation strength is a hyperparameter that can only be compared on the validation set. Too weak and overfitting persists; too strong and the model is crushed into underfitting. The balance must be measured, not derived.

重要公式

L_total = L_data + λ · Σ_w w²
An L2-regularised objective: on top of the data loss, penalise the sum of squared weights, with λ controlling the strength — the weight-decay coefficient.

応用場面

  • Deep learning: Dropout, weight decay and early stopping are near-standard
  • Small-data modelling: fewer samples mean more overfitting, so regularisation and cross-validation matter most
  • Tree models: limiting depth or minimum leaf size is a direct penalty on complexity
  • Fine-tuning large models: low learning rates and light regularisation prevent catastrophic forgetting

よくある誤解

  • Low training loss is not good news. It may simply be a symptom of overfitting; the only trustworthy signal is performance on the validation or test set.
  • Regularisation is not free. Too much of it underfits the model and erases useful signal along with the noise; it must match the data volume and task complexity.
  • Data volume matters more than regularisation. Adding real, diverse samples usually beats any regularisation trick; regularisation is merely compensation for insufficient data.

重要用語

Generalisation
Performance on data the model has not seen
Weight decay (L2)
Adding a squared-weight penalty to the loss to suppress large weights
Dropout
Randomly silencing units during training to prevent co-adaptation
Early stopping
Halting training before validation loss turns upward

参考文献