Classification d’images
Décider à quelle catégorie appartient toute l’image
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
Comment c'est fait
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
Produits représentatifs
4GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
Hunyuan
2023La famille de modèles généraux de Tencent, avec des versions à poids ouverts
Organisations concernées
Usages typiques
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
Comment on l'évalue
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
Limites et points difficiles
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes
Concepts sous-jacents
Classification d’images
Ne lui dites pas « les chats ont des moustaches » : montrez-lui assez de chats
Réseaux de neurones convolutifs
Remplacer les connexions denses par « regarder localement et réutiliser le même filtre partout » : l’idée qui a rendu la reconnaissance d’images enfin fonctionnelle
Apprentissage supervisé
Des paires question-réponse apprennent au modèle à répondre seul