Saltar al contenido
Atlas de IA

Compromiso sesgo-varianza

Todo error de predicción se divide en tres: un modelo demasiado simple, demasiado inestable, y el azar del propio mundo

02 Aprendizaje automáticoExpertoEntrada 6 de este dominio

El texto completo se presenta en inglés; el título y el resumen están traducidos.

DEFINICIÓN

The bias–variance decomposition splits a model’s expected prediction error into three parts: bias² measures how far the model’s average prediction drifts from the true regularity (the source of underfitting); variance measures how sensitive the model is to perturbations of the training set (the source of overfitting); and irreducible error is the floor set by noise in the data that no model can remove. Classical theory holds that the first two trade off, producing a U-shaped error curve.

Intuición

Picture a squad of archers aiming at one bullseye. Bias is how far the whole group of holes sits from the centre — aim wrongly and even a steady hand is useless; variance is how scattered the holes are from one another — a shaky hand misses even with perfect aim on average. The target itself also wobbles in a way no archer can remove (irreducible error). A good model aims true and holds steady, yet the two are often at odds.

Fig. 1

Decomposition of expected generalisation error: bias², variance and irreducible noise each hold a share, and only their sum is the total

Expected generalisation errorExpected generalisation errorBias²Bias²Model too simpleModel too simpleVarianceVarianceModel too complexModel too complexIrreducible error σ²Irreducible error σ²Noise in the data itselfNoise in the data itself
Fig. 2

The double-descent phenomenon (schematic): as model complexity grows, test error falls, rises, then falls again past the interpolation threshold — the classical U-curve breaks

Cómo funciona

  1. 01

    Bias: the model’s systematic drift

    Train the same model family on different training sets and see how far their average prediction sits from the true function. When the model is too simple — a straight line through a curve — bias is large and underfitting follows.

  2. 02

    Variance: the model’s jitter

    Look at how scattered those predictions are from one another. When the model is too complex, swapping a few training samples changes the learned function drastically — high variance, hence overfitting.

  3. 03

    Irreducible error: the data’s floor

    Even with the true function in hand, as long as labels carry noise the error will not fall to zero. This is the problem’s intrinsic floor, independent of the model and impossible to tune away.

  4. 04

    Trade-off and its counterexample

    The classical conclusion is that rising complexity lowers bias and raises variance, so total error follows a U. Yet modern over-parameterised models — with far more parameters than samples — exhibit “double descent”, where error falls, rises, then falls again beyond the interpolation threshold.

Fórmula clave

E[(y − f̂(x))²] = Bias²(f̂) + Var(f̂) + σ²
The three-term decomposition of expected generalisation error: bias² plus variance plus irreducible noise — their sum is the total error.

Dónde se usa

  • Model selection: explains why moderate complexity often beats both extremes
  • Ensembles: bagging mainly cuts variance and boosting mainly cuts bias, which explains random forests versus gradient boosting
  • Regularisation design: weight decay and dropout trade a rise in bias for a fall in variance
  • Modern perspective: explains why very large models still generalise, prompting revisions to classical theory

Errores comunes

  • High variance is not always bad. Double descent shows that in the over-parameterised regime, more capacity can actually lower test error — the classical U-curve is not a universal law.
  • Bias and variance are an expectation-level decomposition. Neither can be measured directly for a single model from a single training run; it is an explanatory framework rather than an operational formula.
  • “More complex means worse” is a misreading. Modern practice repeatedly shows that, given enough data and compute, larger and more over-parameterised models often generalise better.

Términos clave

Bias
How far the model’s average prediction departs from the true regularity
Variance
How sensitive the model is to perturbations of the training set
Irreducible error
The unavoidable error floor caused by label noise
Double descent
The modern counterexample where test error falls again past the interpolation point

Lecturas complementarias