Aller au contenu
Atlas de l'IA

Réseaux de neurones convolutifs

Remplacer les connexions denses par « regarder localement et réutiliser le même filtre partout » : l’idée qui a rendu la reconnaissance d’images enfin fonctionnelle

03 Apprentissage profondIntermédiaireEntrée 4 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

A convolutional neural network is designed for grid-like data such as images. It slides a set of small filters across the input, performing the same weighted sum at every position to extract local features, and composes simple features (edges, corners) into complex ones (textures, parts, objects) layer by layer. Because one filter is reused across the whole image, the parameter count is independent of image size and translation equivariance comes for free.

Intuition

To find a pattern in a large jigsaw, you do not need to view it all at once; you slide a small magnifier cell by cell, noting "right angle" or "arc", then use a bigger magnifier to assemble those marks into an "eye" or a "wheel". The key insight: the same pattern is recognised by the same ruler whether it appears top-left or bottom-right — that is weight sharing, and the fundamental reason a convolution uses fewer parameters than a dense layer.

Fig. 1

A typical convolutional stack: starting from raw pixels, convolution and pooling alternate, the receptive field grows layer by layer, and the features are finally aggregated into a class decision

Shallow layers extract edges and textures; deep layers compose them into parts and objects. Depth sets the level of abstraction.
Fig. 2

ImageNet top-5 error (lower is better): from 16.4% down to the 2% range, the real ten-year progression of convolutional architectures, with human performance around 5.1% for reference

Fonctionnement

  1. 01

    Local receptive field: each output sees only a patch

    Each output connects to a local window of the input rather than all pixels. Adjacent pixels in an image are highly correlated, so local connections suffice to capture edges and corners while cutting the number of connections sharply.

  2. 02

    Weight sharing: one filter slides over the whole image

    The same filter is reused across the entire image, so the parameter count is independent of image size. This both shrinks the model and yields translation equivariance — a cat shifted a few pixels is still a cat.

  3. 03

    Pooling: shrink size, enlarge the receptive field

    Pooling such as max pooling takes the strongest response in a local window, shrinking spatial size while preserving salient activations. It enlarges the effective receptive field of later layers and adds robustness to small shifts.

  4. 04

    Stacking: from edges to objects

    Stacking convolution layers grows the receptive field layer by layer — shallow layers see edges, deep layers see a whole face. A final dense layer or global pooling aggregates the features into a class decision.

Fig. 3

The evolution of convolutional architectures: each generation answers the same question — how to make the network deeper and use the spatial structure of images more effectively

1998LeNet-5The first mature example of conv + pooling + dense, for handwritten digits2012AlexNetReLU + GPU + Dropout; 16.4% ImageNet top-5 error2014VGG / GoogLeNetUniform 3×3 stacks and the Inception module; error down to 6.7%2015ResNetResidual connections make 152 layers trainable; 3.57% error2019EfficientNetCompound scaling jointly tunes depth, width and resolution

Où c'est utilisé

  • Image classification: from handwritten digits to the millions of images in ImageNet
  • Backbones for object detection and semantic segmentation, supplying hierarchical features
  • Defect detection in medical imaging, remote sensing and industrial inspection
  • One-dimensional convolutions are also used for spectrograms and sensor signals

Idées fausses courantes

  • A convolution’s translation invariance is approximate. Large rotations or rescalings still break it, usually addressed with data augmentation or specialised architectures.
  • Depth is not a goal in itself: before residual connections, simply going deeper raised the training error (the degradation problem), which is not overfitting.
  • A small parameter count does not mean a small compute cost. Convolutions are expensive at high spatial resolution in early layers, which is exactly where inference optimisation focuses.

Termes clés

Kernel / filter
A set of learnable weights that slides over the input
Receptive field
The region of the original input that a given output covers
Weight sharing
Reusing one set of weights across all spatial positions
Equivariance
When the input shifts, the output shifts accordingly rather than changing

Lectures complémentaires