Détection d’anomalies
Repérer les quelques cas étranges parmi une masse de données normales
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes data records or time series and outputs how far each deviates from the norm, or a flag. It usually works with few or no labelled anomalies, modelling what normal looks like and calling deviations anomalous. Unlike text classification the anomaly classes are not fixed, and in training you often do not know what an anomaly looks like at all.
Comment c'est fait
Unsupervised routes model normal data: autoencoders flag large reconstruction errors, isolation forests treat easily isolated points as anomalous, and statistical methods fit a distribution and score low-probability points. Time series often use a predictive framing, taking the residual between actual and forecast values as the score. With a few labels, semi-supervised or contrastive learning takes over, and cost-sensitive thresholds balance misses against false alarms.
Produits représentatifs
3NeMo
2019Boîte à outils pour entraîner et personnaliser de grands modèles
Hugging Face Hub
2016Le carrefour des modèles et jeux de données ouverts
Transformers
2018Une API pour charger et entraîner des modèles pré-entraînés
Organisations concernées
Usages typiques
- Early warning for equipment and production-line faults
- Spotting oddity in transactions and money laundering
- Intrusion and abnormal-traffic monitoring
- Quality sampling and data-entry error catching
Comment on l'évalue
- AUROC
- Threshold-free ranking quality, usable on heavily imbalanced data
- PR-AUC
- More informative than AUROC when anomalies are extremely rare
- Detection delay
- Time from onset to alert, key in real-time settings
Limites et points difficiles
- With so few anomalies a small threshold change spikes false alarms, and operators quickly go numb to them
- Concept drift marks normal behaviour as anomalous, and seasonality or business change demands continual recalibration
- In high dimensions with correlated features distance-based measures break down, diluting the gaps between normal points
Concepts sous-jacents
Apprentissage non supervisé
Sans réponses, la structure doit émerger des données elles-mêmes — et il faut redéfinir ce qui est « bon »
Probabilités et distributions
Un modèle ne fournit pas une réponse, mais un degré de croyance pour chaque réponse possible
Évaluation de modèle et validation croisée
L’exactitude est la métrique la plus trompeuse — une évaluation erronée fait tout s’effondrer