Detección de anomalías
Señalar los pocos casos raros entre muchos normales
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes data records or time series and outputs how far each deviates from the norm, or a flag. It usually works with few or no labelled anomalies, modelling what normal looks like and calling deviations anomalous. Unlike text classification the anomaly classes are not fixed, and in training you often do not know what an anomaly looks like at all.
Cómo se consigue técnicamente
Unsupervised routes model normal data: autoencoders flag large reconstruction errors, isolation forests treat easily isolated points as anomalous, and statistical methods fit a distribution and score low-probability points. Time series often use a predictive framing, taking the residual between actual and forecast values as the score. With a few labels, semi-supervised or contrastive learning takes over, and cost-sensitive thresholds balance misses against false alarms.
Productos representativos
3NeMo
2019Kit de herramientas para entrenar y personalizar modelos grandes
Hugging Face Hub
2016El punto de encuentro de modelos y datos abiertos
Transformers
2018Una API para cargar y entrenar modelos preentrenados
Organizaciones relacionadas
Usos típicos
- Early warning for equipment and production-line faults
- Spotting oddity in transactions and money laundering
- Intrusion and abnormal-traffic monitoring
- Quality sampling and data-entry error catching
Cómo se evalúa
- AUROC
- Threshold-free ranking quality, usable on heavily imbalanced data
- PR-AUC
- More informative than AUROC when anomalies are extremely rare
- Detection delay
- Time from onset to alert, key in real-time settings
Límites y dificultades
- With so few anomalies a small threshold change spikes false alarms, and operators quickly go numb to them
- Concept drift marks normal behaviour as anomalous, and seasonality or business change demands continual recalibration
- In high dimensions with correlated features distance-based measures break down, diluting the gaps between normal points
Conceptos detrás
Aprendizaje no supervisado
Sin respuestas, la estructura debe emerger de los propios datos — y hay que redefinir qué es «bueno»
Probabilidad y distribuciones
Un modelo nunca te da una respuesta; te da un grado de creencia sobre cada respuesta posible
Evaluación de modelos y validación cruzada
La exactitud es la métrica más fácil de engañar — si evalúas mal, todo lo demás cae