Skip to content
AI Atlas

Convolution Operations

One small stencil swept across the image finds edges, textures and shapes

05 Computer VisionBeginnerEntry 2 in this domain

DEFINITION

Convolution slides a small matrix (the kernel) across an image, and at each position computes an element-wise product followed by a sum to produce a single output value. Repeating this “local weighted sum” at every location with shared weights lets it detect local patterns such as edges and textures using very few parameters.

Intuition

Picture a small stencil engraved with −1, 0 and +1 pressed against a patch of the photo: if the patch is a vertical edge, the weighted sum through the stencil comes out large; over flat sky it is near zero. Sweep the same stencil across the whole photo and you get a response map of “where edges of this kind are.” What the network learns is precisely the numbers on those stencils.

Fig. 1

The geometry of convolution: a 5×5 patch through a 3×3 kernel (stride 1, no padding) yields a 3×3 response map in which the object boundary is extracted

Fig. 2

Why images use convolution rather than fully connected layers: a comparison along locality, weight sharing and translation equivariance

  • Convolution
  • Fully connected

How it works

  1. 01

    Slide and take the weighted sum

    The kernel covers a small window of the input, multiplies element-wise and sums, producing one value on the output map. A larger kernel enlarges the receptive field but also raises parameters and compute.

  2. 02

    Stride and padding control the output size

    Stride sets how far the window jumps (downsampling); padding pads the border with zeros to preserve size. Together with the kernel size they fix the output dimensions: out = ⌊(in + 2·pad − kernel)/stride⌋ + 1.

  3. 03

    Dilated and transposed convolution

    Dilated convolution inserts gaps between kernel elements, enlarging the receptive field without new parameters and appearing widely in segmentation; transposed convolution goes the other way, upsampling feature maps to restore resolution or generate.

  4. 04

    Stacking layers: from edges to objects

    Shallow kernels learn edges and colour blobs, middle layers compose them into textures and parts, deep layers into objects. The key point: these kernels are not hand-designed — backpropagation learns them from data.

Key formula

out = ⌊(H + 2P − K) / S⌋ + 1
H is the input side length, K the kernel size, P the padding and S the stride. A stride above 1 shrinks the output, i.e. downsamples.
Fig. 3

The weights of an edge-detection kernel: a positive centre surrounded by negative neighbours, highly sensitive to abrupt brightness changes (hover a cell for its weight)

-1-1-1-18-1-1-1-1

Where it is used

  • Feature-extraction backbones: the bottom of nearly every CNN, replacing hand-designed operators such as SIFT and HOG
  • Edge and gradient detection: classic kernels like Sobel and Laplacian remain in use for image processing and industrial inspection
  • Upsampling and generation: transposed convolution serves segmentation, super-resolution and decoders
  • Large-receptive-field modelling: dilated convolution is used in semantic segmentation, audio and sequence modelling

Common misconceptions

  • Convolution is not synonymous with “blurring”. Blur is merely one kernel among many; the same sliding mechanism with different weights sharpens, detects edges or responds to a specific orientation.
  • Fewer parameters does not mean less computation. Weight sharing collapses the parameter count, but every location is still evaluated, so the cost grows with image area.
  • Bigger kernels are not automatically better. Modern designs prefer stacking several 3×3 kernels to reach the same receptive field with fewer parameters and more nonlinearity.

Key terms

Kernel / filter
The small weight matrix that is learned
Stride
How many pixels the window jumps each step
Padding
Adding zeros at the border to control output size
Receptive field
The input region that one output pixel depends on

Further reading