本文へスキップ
AI図鑑

モデル評価と交差検証

正解率は最も人を欺く指標——評価を誤れば、あとはすべて崩れる

02 機械学習中級この領域の第 5 項目

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

Model evaluation measures a model’s true performance using a separate dataset that never took part in training. It involves choosing suitable metrics (accuracy, precision, recall, F1, ROC-AUC), reducing the randomness of the estimate with cross-validation, and guarding against data leakage — any peeking at test information during training will inflate the reported numbers.

直観的な理解

It is like testing a student on an exam they have never seen. If the paper happens to be the very problems they drilled, a high score proves nothing. The subtler trap is a leaked answer key: accidentally slipping the answers into the revision material. Any overreach in splitting, feature engineering, or preprocessing turns evaluation into self-deception.

図 1

A binary confusion matrix (1000 samples, positives only 10%): 94.5% accuracy looks good, but the distribution across the four cells is the real story

85154086
図 2

F1 of several classifiers on one imbalanced dataset: models with similar accuracy can differ widely in F1

仕組み

  1. 01

    Split the data strictly

    Training, validation and test sets must be mutually isolated. Normalisation, feature selection and hyperparameter search may only touch the training portion, or information leaks from test into training.

  2. 02

    Read the four outcomes from the confusion matrix

    True positives, true negatives, false positives, false negatives. Almost every metric is a combination of these four counts; read the matrix first and the metrics will not mislead you.

  3. 03

    Choose metrics by cost

    When missing a positive is costly, watch recall; when false alarms are costly, watch precision; when both matter, use F1 or the PR curve. Under class imbalance, accuracy is almost guaranteed to mislead.

  4. 04

    Cross-validate and quantify uncertainty

    Rotate the data into k folds, training on k−1 and validating on the remaining one, then average. On limited data this yields a steadier estimate and, in passing, tells you how uncertain that estimate is.

重要公式

F1 = 2 · (P · R) / (P + R)
F1 is the harmonic mean of precision P and recall R; it is high only when both are high — neither can be sacrificed.

応用場面

  • Model selection: compare candidate models or hyperparameters fairly
  • Threshold tuning: pick a decision threshold balancing precision and recall via ROC/PR curves
  • Pre-launch acceptance: simulate deployment with a held-out test set or a time-based split
  • Competitions and benchmarks: a shared evaluation protocol makes different methods comparable

よくある誤解

  • The accuracy trap: on data where positives are only 1%, predicting “negative” for everything scores 99% accuracy and is utterly useless.
  • Data leakage: randomly splitting time-ordered data lets future information seep into training; standardising with statistics computed over the full dataset is another common source.
  • Misusing cross-validation: random k-fold on time series overestimates performance, and tuning repeatedly until the test set looks good is the same as training on the test set.

重要用語

Confusion matrix
A cross-tabulation of true versus predicted classes
Precision & recall
Precision asks how many alerts are real; recall asks how many real cases were caught
F1
The harmonic mean of precision and recall
ROC-AUC
Area under the ROC curve, measuring ranking ability across all thresholds

参考文献