Ir para o conteúdo
Atlas de IA

Redes neurais convolucionais

Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional

03 Aprendizado profundoIntermediárioEntrada 4 deste domínio

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

A convolutional neural network is designed for grid-like data such as images. It slides a set of small filters across the input, performing the same weighted sum at every position to extract local features, and composes simple features (edges, corners) into complex ones (textures, parts, objects) layer by layer. Because one filter is reused across the whole image, the parameter count is independent of image size and translation equivariance comes for free.

Intuição

To find a pattern in a large jigsaw, you do not need to view it all at once; you slide a small magnifier cell by cell, noting "right angle" or "arc", then use a bigger magnifier to assemble those marks into an "eye" or a "wheel". The key insight: the same pattern is recognised by the same ruler whether it appears top-left or bottom-right — that is weight sharing, and the fundamental reason a convolution uses fewer parameters than a dense layer.

Fig. 1

A typical convolutional stack: starting from raw pixels, convolution and pooling alternate, the receptive field grows layer by layer, and the features are finally aggregated into a class decision

Shallow layers extract edges and textures; deep layers compose them into parts and objects. Depth sets the level of abstraction.
Fig. 2

ImageNet top-5 error (lower is better): from 16.4% down to the 2% range, the real ten-year progression of convolutional architectures, with human performance around 5.1% for reference

Como funciona

  1. 01

    Local receptive field: each output sees only a patch

    Each output connects to a local window of the input rather than all pixels. Adjacent pixels in an image are highly correlated, so local connections suffice to capture edges and corners while cutting the number of connections sharply.

  2. 02

    Weight sharing: one filter slides over the whole image

    The same filter is reused across the entire image, so the parameter count is independent of image size. This both shrinks the model and yields translation equivariance — a cat shifted a few pixels is still a cat.

  3. 03

    Pooling: shrink size, enlarge the receptive field

    Pooling such as max pooling takes the strongest response in a local window, shrinking spatial size while preserving salient activations. It enlarges the effective receptive field of later layers and adds robustness to small shifts.

  4. 04

    Stacking: from edges to objects

    Stacking convolution layers grows the receptive field layer by layer — shallow layers see edges, deep layers see a whole face. A final dense layer or global pooling aggregates the features into a class decision.

Fig. 3

The evolution of convolutional architectures: each generation answers the same question — how to make the network deeper and use the spatial structure of images more effectively

1998LeNet-5The first mature example of conv + pooling + dense, for handwritten digits2012AlexNetReLU + GPU + Dropout; 16.4% ImageNet top-5 error2014VGG / GoogLeNetUniform 3×3 stacks and the Inception module; error down to 6.7%2015ResNetResidual connections make 152 layers trainable; 3.57% error2019EfficientNetCompound scaling jointly tunes depth, width and resolution

Onde é usado

  • Image classification: from handwritten digits to the millions of images in ImageNet
  • Backbones for object detection and semantic segmentation, supplying hierarchical features
  • Defect detection in medical imaging, remote sensing and industrial inspection
  • One-dimensional convolutions are also used for spectrograms and sensor signals

Equívocos comuns

  • A convolution’s translation invariance is approximate. Large rotations or rescalings still break it, usually addressed with data augmentation or specialised architectures.
  • Depth is not a goal in itself: before residual connections, simply going deeper raised the training error (the degradation problem), which is not overfitting.
  • A small parameter count does not mean a small compute cost. Convolutions are expensive at high spatial resolution in early layers, which is exactly where inference optimisation focuses.

Termos-chave

Kernel / filter
A set of learnable weights that slides over the input
Receptive field
The region of the original input that a given output covers
Weight sharing
Reusing one set of weights across all spatial positions
Equivariance
When the input shifts, the output shifts accordingly rather than changing

Leituras complementares