Operações de convolução
Um pequeno molde varrendo a imagem encontra bordas, texturas e formas
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
DEFINIÇÃO
Convolution slides a small matrix (the kernel) across an image, and at each position computes an element-wise product followed by a sum to produce a single output value. Repeating this “local weighted sum” at every location with shared weights lets it detect local patterns such as edges and textures using very few parameters.
Intuição
Picture a small stencil engraved with −1, 0 and +1 pressed against a patch of the photo: if the patch is a vertical edge, the weighted sum through the stencil comes out large; over flat sky it is near zero. Sweep the same stencil across the whole photo and you get a response map of “where edges of this kind are.” What the network learns is precisely the numbers on those stencils.
The geometry of convolution: a 5×5 patch through a 3×3 kernel (stride 1, no padding) yields a 3×3 response map in which the object boundary is extracted
Why images use convolution rather than fully connected layers: a comparison along locality, weight sharing and translation equivariance
- Convolution
- Fully connected
Como funciona
- 01
Slide and take the weighted sum
The kernel covers a small window of the input, multiplies element-wise and sums, producing one value on the output map. A larger kernel enlarges the receptive field but also raises parameters and compute.
- 02
Stride and padding control the output size
Stride sets how far the window jumps (downsampling); padding pads the border with zeros to preserve size. Together with the kernel size they fix the output dimensions: out = ⌊(in + 2·pad − kernel)/stride⌋ + 1.
- 03
Dilated and transposed convolution
Dilated convolution inserts gaps between kernel elements, enlarging the receptive field without new parameters and appearing widely in segmentation; transposed convolution goes the other way, upsampling feature maps to restore resolution or generate.
- 04
Stacking layers: from edges to objects
Shallow kernels learn edges and colour blobs, middle layers compose them into textures and parts, deep layers into objects. The key point: these kernels are not hand-designed — backpropagation learns them from data.
Fórmula-chave
out = ⌊(H + 2P − K) / S⌋ + 1The weights of an edge-detection kernel: a positive centre surrounded by negative neighbours, highly sensitive to abrupt brightness changes (hover a cell for its weight)
Onde é usado
- Feature-extraction backbones: the bottom of nearly every CNN, replacing hand-designed operators such as SIFT and HOG
- Edge and gradient detection: classic kernels like Sobel and Laplacian remain in use for image processing and industrial inspection
- Upsampling and generation: transposed convolution serves segmentation, super-resolution and decoders
- Large-receptive-field modelling: dilated convolution is used in semantic segmentation, audio and sequence modelling
Equívocos comuns
- Convolution is not synonymous with “blurring”. Blur is merely one kernel among many; the same sliding mechanism with different weights sharpens, detects edges or responds to a specific orientation.
- Fewer parameters does not mean less computation. Weight sharing collapses the parameter count, but every location is still evaluated, so the cost grows with image area.
- Bigger kernels are not automatically better. Modern designs prefer stacking several 3×3 kernels to reach the same receptive field with fewer parameters and more nonlinearity.
Termos-chave
- Kernel / filter
- The small weight matrix that is learned
- Stride
- How many pixels the window jumps each step
- Padding
- Adding zeros at the border to control output size
- Receptive field
- The input region that one output pixel depends on