Image Classification
Decide what category the whole image is
WHAT THIS CAPABILITY MEANS
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
How it is done
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
Representative products
4GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Organizations involved
Typical uses
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
How it is evaluated
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
Limits and hard parts
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes
Concepts behind it
Image Classification
Don’t tell it “cats have whiskers” — show it enough cats and it works it out
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Supervised Learning
Pairs of questions and answers teach a model to answer on its own