Evaluación de modelos y validación cruzada
La exactitud es la métrica más fácil de engañar — si evalúas mal, todo lo demás cae
El texto completo se presenta en inglés; el título y el resumen están traducidos.
DEFINICIÓN
Model evaluation measures a model’s true performance using a separate dataset that never took part in training. It involves choosing suitable metrics (accuracy, precision, recall, F1, ROC-AUC), reducing the randomness of the estimate with cross-validation, and guarding against data leakage — any peeking at test information during training will inflate the reported numbers.
Intuición
It is like testing a student on an exam they have never seen. If the paper happens to be the very problems they drilled, a high score proves nothing. The subtler trap is a leaked answer key: accidentally slipping the answers into the revision material. Any overreach in splitting, feature engineering, or preprocessing turns evaluation into self-deception.
A binary confusion matrix (1000 samples, positives only 10%): 94.5% accuracy looks good, but the distribution across the four cells is the real story
F1 of several classifiers on one imbalanced dataset: models with similar accuracy can differ widely in F1
Cómo funciona
- 01
Split the data strictly
Training, validation and test sets must be mutually isolated. Normalisation, feature selection and hyperparameter search may only touch the training portion, or information leaks from test into training.
- 02
Read the four outcomes from the confusion matrix
True positives, true negatives, false positives, false negatives. Almost every metric is a combination of these four counts; read the matrix first and the metrics will not mislead you.
- 03
Choose metrics by cost
When missing a positive is costly, watch recall; when false alarms are costly, watch precision; when both matter, use F1 or the PR curve. Under class imbalance, accuracy is almost guaranteed to mislead.
- 04
Cross-validate and quantify uncertainty
Rotate the data into k folds, training on k−1 and validating on the remaining one, then average. On limited data this yields a steadier estimate and, in passing, tells you how uncertain that estimate is.
Fórmula clave
F1 = 2 · (P · R) / (P + R)Dónde se usa
- Model selection: compare candidate models or hyperparameters fairly
- Threshold tuning: pick a decision threshold balancing precision and recall via ROC/PR curves
- Pre-launch acceptance: simulate deployment with a held-out test set or a time-based split
- Competitions and benchmarks: a shared evaluation protocol makes different methods comparable
Errores comunes
- The accuracy trap: on data where positives are only 1%, predicting “negative” for everything scores 99% accuracy and is utterly useless.
- Data leakage: randomly splitting time-ordered data lets future information seep into training; standardising with statistics computed over the full dataset is another common source.
- Misusing cross-validation: random k-fold on time series overestimates performance, and tuning repeatedly until the test set looks good is the same as training on the test set.
Términos clave
- Confusion matrix
- A cross-tabulation of true versus predicted classes
- Precision & recall
- Precision asks how many alerts are real; recall asks how many real cases were caught
- F1
- The harmonic mean of precision and recall
- ROC-AUC
- Area under the ROC curve, measuring ranking ability across all thresholds