본문으로 건너뛰기
AI 도감

이상 탐지

대량의 정상 데이터에서 이상한 소수를 찾는다

데이터와 문서중급 #40
입력표텍스트

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes data records or time series and outputs how far each deviates from the norm, or a flag. It usually works with few or no labelled anomalies, modelling what normal looks like and calling deviations anomalous. Unlike text classification the anomaly classes are not fixed, and in training you often do not know what an anomaly looks like at all.

기술적으로 구현하는 방법

Unsupervised routes model normal data: autoencoders flag large reconstruction errors, isolation forests treat easily isolated points as anomalous, and statistical methods fit a distribution and score low-probability points. Time series often use a predictive framing, taking the residual between actual and forecast values as the score. With a few labels, semi-supervised or contrastive learning takes over, and cost-sensitive thresholds balance misses against false alarms.

대표 제품

3

관련 기관

대표적 용도

  • Early warning for equipment and production-line faults
  • Spotting oddity in transactions and money laundering
  • Intrusion and abnormal-traffic monitoring
  • Quality sampling and data-entry error catching

성능을 평가하는 방법

AUROC
Threshold-free ranking quality, usable on heavily imbalanced data
PR-AUC
More informative than AUROC when anomalies are extremely rare
Detection delay
Time from onset to alert, key in real-time settings

경계와 난점

  • With so few anomalies a small threshold change spikes false alarms, and operators quickly go numb to them
  • Concept drift marks normal behaviour as anomalous, and seasonality or business change demands continual recalibration
  • In high dimensions with correlated features distance-based measures break down, diluting the gaps between normal points

뒤에 있는 개념