Phân loại ảnh
Đừng nói “mèo có râu” — cho nó xem đủ nhiều mèo, nó tự rút ra
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
ĐỊNH NGHĨA
Image classification takes an image and outputs which of a predefined set of classes it belongs to, usually as a probability per class. It says nothing about where objects are or how many there are — only “what this image is overall.” It is the most fundamental vision task and the earliest point at which deep learning surged.
Trực giác
Think of a classifier as a well-travelled appraiser. You hand over a photo; in his head he compares it against the tens of millions he has seen, then reports “85% cat, 10% dog, 3% fox.” What deep learning does is train that comparative instinct into a computable procedure, using a huge pile of labelled images.
A decade on ImageNet: from AlexNet to ViT, error falls off a cliff — with data scale as the ever-present prerequisite
Top-5 error of the ImageNet winners falls year by year, crossing human level around 2015 (the dashed line marks the ~5.1% human reference)
- ILSVRC winner
- Human level (~5.1%)
Cách hoạt động
- 01
Feature extraction → classification head
A convolutional backbone progressively compresses the image into a high-level feature vector; a final fully connected (or linear) layer maps it to K class scores, and softmax normalises them into a probability distribution.
- 02
Loss and backpropagation
Cross-entropy measures the gap between the predicted distribution and the true label; the gradient flows back and updates both backbone and head, all parameters moved by a single loss at once.
- 03
ImageNet and architectures co-evolving
In 2012 AlexNet used an 8-layer convolutional network, two GPUs and 1.2 million images to drop top-5 error from 25.8% to 15.3%. Then VGG, GoogLeNet and ResNet pushed the error down close to human level.
- 04
ViT: treating an image as a sequence of words
ViT (2020) cuts the image into 16×16 patches, linearly embeds them and feeds them to a Transformer as tokens. It trails CNNs on modest data but overtakes them given hundreds of millions of images such as JFT-300M — showing that an architecture’s advantage depends on data scale.
The CNN-versus-ViT trade-off: which relies more on data and which on inductive bias
Ứng dụng
- Content moderation and photo organisation: automatic tagging and retrieval
- Multi-class medical triage: classifying X-rays and pathology slides
- Industrial inspection: pass/fail judgement and defect categorisation
- As a pre-training task: transferring learned features to detection, segmentation and beyond
Hiểu lầm thường gặp
- The single-label assumption: an image often contains several objects, and ImageNet-style single-label training hides this, forcing the model to “pick one” on multi-object images.
- Class imbalance and long tails: rare classes perform poorly, yet overall accuracy looks fine, masking the real weakness.
- “Superhuman” deserves caution: human top-5 error on ImageNet is about 5.1%, but that is a comparison within a fixed label set and annotation scheme — not evidence of general visual understanding.
Thuật ngữ chính
- Top-1 / top-5 error
- Whether the top prediction / top five include the true label
- Cross-entropy loss
- The standard objective for classification training
- Transfer learning
- Pre-train on a large dataset, then fine-tune on a small task
- ViT
- An architecture that applies a Transformer to image patches