Перейти к содержимому
Атлас ИИ

Классификация изображений

Не говорите «у кошек есть усы» — покажите достаточно кошек, и он поймёт сам

05 Компьютерное зрениеСреднийСтатья 3 в этой области

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ОПРЕДЕЛЕНИЕ

Image classification takes an image and outputs which of a predefined set of classes it belongs to, usually as a probability per class. It says nothing about where objects are or how many there are — only “what this image is overall.” It is the most fundamental vision task and the earliest point at which deep learning surged.

Интуиция

Think of a classifier as a well-travelled appraiser. You hand over a photo; in his head he compares it against the tens of millions he has seen, then reports “85% cat, 10% dog, 3% fox.” What deep learning does is train that comparative instinct into a computable procedure, using a huge pile of labelled images.

Рис. 1

A decade on ImageNet: from AlexNet to ViT, error falls off a cliff — with data scale as the ever-present prerequisite

2012AlexNet8 conv layers on two GPUs; 15.3% top-5 error2014VGG / GoogLeNetDeeper networks and the Inception module; error down to 6.7%2015ResNetResidual connections make 152 layers trainable; 3.57% error2017SENetChannel attention; 2.25% error with ensembles2020ViTImage patches + Transformer; needs massive data to win
Рис. 2

Top-5 error of the ImageNet winners falls year by year, crossing human level around 2015 (the dashed line marks the ~5.1% human reference)

  • ILSVRC winner
  • Human level (~5.1%)

Как это работает

  1. 01

    Feature extraction → classification head

    A convolutional backbone progressively compresses the image into a high-level feature vector; a final fully connected (or linear) layer maps it to K class scores, and softmax normalises them into a probability distribution.

  2. 02

    Loss and backpropagation

    Cross-entropy measures the gap between the predicted distribution and the true label; the gradient flows back and updates both backbone and head, all parameters moved by a single loss at once.

  3. 03

    ImageNet and architectures co-evolving

    In 2012 AlexNet used an 8-layer convolutional network, two GPUs and 1.2 million images to drop top-5 error from 25.8% to 15.3%. Then VGG, GoogLeNet and ResNet pushed the error down close to human level.

  4. 04

    ViT: treating an image as a sequence of words

    ViT (2020) cuts the image into 16×16 patches, linearly embeds them and feeds them to a Transformer as tokens. It trails CNNs on modest data but overtakes them given hundreds of millions of images such as JFT-300M — showing that an architecture’s advantage depends on data scale.

Рис. 3

The CNN-versus-ViT trade-off: which relies more on data and which on inductive bias

Области применения

  • Content moderation and photo organisation: automatic tagging and retrieval
  • Multi-class medical triage: classifying X-rays and pathology slides
  • Industrial inspection: pass/fail judgement and defect categorisation
  • As a pre-training task: transferring learned features to detection, segmentation and beyond

Частые заблуждения

  • The single-label assumption: an image often contains several objects, and ImageNet-style single-label training hides this, forcing the model to “pick one” on multi-object images.
  • Class imbalance and long tails: rare classes perform poorly, yet overall accuracy looks fine, masking the real weakness.
  • “Superhuman” deserves caution: human top-5 error on ImageNet is about 5.1%, but that is a comparison within a fixed label set and annotation scheme — not evidence of general visual understanding.

Ключевые термины

Top-1 / top-5 error
Whether the top prediction / top five include the true label
Cross-entropy loss
The standard objective for classification training
Transfer learning
Pre-train on a large dataset, then fine-tune on a small task
ViT
An architecture that applies a Transformer to image patches

Дополнительная литература