Ir para o conteúdo
Atlas de IA

Avaliação de modelos e validação cruzada

A acurácia é a métrica mais fácil de enganar — erre na avaliação e tudo o mais desmorona

02 Aprendizado de máquinaIntermediárioEntrada 5 deste domínio

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

Model evaluation measures a model’s true performance using a separate dataset that never took part in training. It involves choosing suitable metrics (accuracy, precision, recall, F1, ROC-AUC), reducing the randomness of the estimate with cross-validation, and guarding against data leakage — any peeking at test information during training will inflate the reported numbers.

Intuição

It is like testing a student on an exam they have never seen. If the paper happens to be the very problems they drilled, a high score proves nothing. The subtler trap is a leaked answer key: accidentally slipping the answers into the revision material. Any overreach in splitting, feature engineering, or preprocessing turns evaluation into self-deception.

Fig. 1

A binary confusion matrix (1000 samples, positives only 10%): 94.5% accuracy looks good, but the distribution across the four cells is the real story

85154086
Fig. 2

F1 of several classifiers on one imbalanced dataset: models with similar accuracy can differ widely in F1

Como funciona

  1. 01

    Split the data strictly

    Training, validation and test sets must be mutually isolated. Normalisation, feature selection and hyperparameter search may only touch the training portion, or information leaks from test into training.

  2. 02

    Read the four outcomes from the confusion matrix

    True positives, true negatives, false positives, false negatives. Almost every metric is a combination of these four counts; read the matrix first and the metrics will not mislead you.

  3. 03

    Choose metrics by cost

    When missing a positive is costly, watch recall; when false alarms are costly, watch precision; when both matter, use F1 or the PR curve. Under class imbalance, accuracy is almost guaranteed to mislead.

  4. 04

    Cross-validate and quantify uncertainty

    Rotate the data into k folds, training on k−1 and validating on the remaining one, then average. On limited data this yields a steadier estimate and, in passing, tells you how uncertain that estimate is.

Fórmula-chave

F1 = 2 · (P · R) / (P + R)
F1 is the harmonic mean of precision P and recall R; it is high only when both are high — neither can be sacrificed.

Onde é usado

  • Model selection: compare candidate models or hyperparameters fairly
  • Threshold tuning: pick a decision threshold balancing precision and recall via ROC/PR curves
  • Pre-launch acceptance: simulate deployment with a held-out test set or a time-based split
  • Competitions and benchmarks: a shared evaluation protocol makes different methods comparable

Equívocos comuns

  • The accuracy trap: on data where positives are only 1%, predicting “negative” for everything scores 99% accuracy and is utterly useless.
  • Data leakage: randomly splitting time-ordered data lets future information seep into training; standardising with statistics computed over the full dataset is another common source.
  • Misusing cross-validation: random k-fold on time series overestimates performance, and tuning repeatedly until the test set looks good is the same as training on the test set.

Termos-chave

Confusion matrix
A cross-tabulation of true versus predicted classes
Precision & recall
Precision asks how many alerts are real; recall asks how many real cases were caught
F1
The harmonic mean of precision and recall
ROC-AUC
Area under the ROC curve, measuring ranking ability across all thresholds

Leituras complementares