Compromisso viés-variância
Todo erro de previsão se divide em três: modelo simples demais, modelo instável demais e a aleatoriedade do próprio mundo
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
DEFINIÇÃO
The bias–variance decomposition splits a model’s expected prediction error into three parts: bias² measures how far the model’s average prediction drifts from the true regularity (the source of underfitting); variance measures how sensitive the model is to perturbations of the training set (the source of overfitting); and irreducible error is the floor set by noise in the data that no model can remove. Classical theory holds that the first two trade off, producing a U-shaped error curve.
Intuição
Picture a squad of archers aiming at one bullseye. Bias is how far the whole group of holes sits from the centre — aim wrongly and even a steady hand is useless; variance is how scattered the holes are from one another — a shaky hand misses even with perfect aim on average. The target itself also wobbles in a way no archer can remove (irreducible error). A good model aims true and holds steady, yet the two are often at odds.
Decomposition of expected generalisation error: bias², variance and irreducible noise each hold a share, and only their sum is the total
The double-descent phenomenon (schematic): as model complexity grows, test error falls, rises, then falls again past the interpolation threshold — the classical U-curve breaks
Como funciona
- 01
Bias: the model’s systematic drift
Train the same model family on different training sets and see how far their average prediction sits from the true function. When the model is too simple — a straight line through a curve — bias is large and underfitting follows.
- 02
Variance: the model’s jitter
Look at how scattered those predictions are from one another. When the model is too complex, swapping a few training samples changes the learned function drastically — high variance, hence overfitting.
- 03
Irreducible error: the data’s floor
Even with the true function in hand, as long as labels carry noise the error will not fall to zero. This is the problem’s intrinsic floor, independent of the model and impossible to tune away.
- 04
Trade-off and its counterexample
The classical conclusion is that rising complexity lowers bias and raises variance, so total error follows a U. Yet modern over-parameterised models — with far more parameters than samples — exhibit “double descent”, where error falls, rises, then falls again beyond the interpolation threshold.
Fórmula-chave
E[(y − f̂(x))²] = Bias²(f̂) + Var(f̂) + σ²Onde é usado
- Model selection: explains why moderate complexity often beats both extremes
- Ensembles: bagging mainly cuts variance and boosting mainly cuts bias, which explains random forests versus gradient boosting
- Regularisation design: weight decay and dropout trade a rise in bias for a fall in variance
- Modern perspective: explains why very large models still generalise, prompting revisions to classical theory
Equívocos comuns
- High variance is not always bad. Double descent shows that in the over-parameterised regime, more capacity can actually lower test error — the classical U-curve is not a universal law.
- Bias and variance are an expectation-level decomposition. Neither can be measured directly for a single model from a single training run; it is an explanatory framework rather than an operational formula.
- “More complex means worse” is a misreading. Modern practice repeatedly shows that, given enough data and compute, larger and more over-parameterised models often generalise better.
Termos-chave
- Bias
- How far the model’s average prediction departs from the true regularity
- Variance
- How sensitive the model is to perturbations of the training set
- Irreducible error
- The unavoidable error floor caused by label noise
- Double descent
- The modern counterexample where test error falls again past the interpolation point