Phát hiện bất thường
Tìm ra vài trường hợp bất thường giữa rất nhiều dữ liệu bình thường
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes data records or time series and outputs how far each deviates from the norm, or a flag. It usually works with few or no labelled anomalies, modelling what normal looks like and calling deviations anomalous. Unlike text classification the anomaly classes are not fixed, and in training you often do not know what an anomaly looks like at all.
Làm ra sao về mặt kỹ thuật
Unsupervised routes model normal data: autoencoders flag large reconstruction errors, isolation forests treat easily isolated points as anomalous, and statistical methods fit a distribution and score low-probability points. Time series often use a predictive framing, taking the residual between actual and forecast values as the score. With a few labels, semi-supervised or contrastive learning takes over, and cost-sensitive thresholds balance misses against false alarms.
Sản phẩm tiêu biểu
3Tổ chức liên quan
Cách dùng tiêu biểu
- Early warning for equipment and production-line faults
- Spotting oddity in transactions and money laundering
- Intrusion and abnormal-traffic monitoring
- Quality sampling and data-entry error catching
Đánh giá nó tốt hay không thế nào
- AUROC
- Threshold-free ranking quality, usable on heavily imbalanced data
- PR-AUC
- More informative than AUROC when anomalies are extremely rare
- Detection delay
- Time from onset to alert, key in real-time settings
Ranh giới và điểm khó
- With so few anomalies a small threshold change spikes false alarms, and operators quickly go numb to them
- Concept drift marks normal behaviour as anomalous, and seasonality or business change demands continual recalibration
- In high dimensions with correlated features distance-based measures break down, diluting the gaps between normal points
Các khái niệm đằng sau
Học không giám sát
Khi không có đáp án, cấu trúc phải tự nổi lên từ dữ liệu — và tiêu chuẩn “tốt” cũng phải định nghĩa lại
Xác suất và phân phối xác suất
Mô hình không đưa ra ‘đáp án’, mà cho biết mức độ tin cậy cho từng khả năng
Đánh giá mô hình và kiểm định chéo
Độ chính xác là chỉ số dễ đánh lừa nhất — đánh giá sai thì mọi thứ sau đó đều vô nghĩa