본문으로 건너뛰기
AI 도감

합성곱 연산

작은 틀 하나를 이미지 전체에 훑려 모서리·질감·형태를 찾아낸다

05 컴퓨터 비전입문이 영역의 2번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Convolution slides a small matrix (the kernel) across an image, and at each position computes an element-wise product followed by a sum to produce a single output value. Repeating this “local weighted sum” at every location with shared weights lets it detect local patterns such as edges and textures using very few parameters.

직관적 이해

Picture a small stencil engraved with −1, 0 and +1 pressed against a patch of the photo: if the patch is a vertical edge, the weighted sum through the stencil comes out large; over flat sky it is near zero. Sweep the same stencil across the whole photo and you get a response map of “where edges of this kind are.” What the network learns is precisely the numbers on those stencils.

그림 1

The geometry of convolution: a 5×5 patch through a 3×3 kernel (stride 1, no padding) yields a 3×3 response map in which the object boundary is extracted

그림 2

Why images use convolution rather than fully connected layers: a comparison along locality, weight sharing and translation equivariance

  • Convolution
  • Fully connected

작동 원리

  1. 01

    Slide and take the weighted sum

    The kernel covers a small window of the input, multiplies element-wise and sums, producing one value on the output map. A larger kernel enlarges the receptive field but also raises parameters and compute.

  2. 02

    Stride and padding control the output size

    Stride sets how far the window jumps (downsampling); padding pads the border with zeros to preserve size. Together with the kernel size they fix the output dimensions: out = ⌊(in + 2·pad − kernel)/stride⌋ + 1.

  3. 03

    Dilated and transposed convolution

    Dilated convolution inserts gaps between kernel elements, enlarging the receptive field without new parameters and appearing widely in segmentation; transposed convolution goes the other way, upsampling feature maps to restore resolution or generate.

  4. 04

    Stacking layers: from edges to objects

    Shallow kernels learn edges and colour blobs, middle layers compose them into textures and parts, deep layers into objects. The key point: these kernels are not hand-designed — backpropagation learns them from data.

핵심 수식

out = ⌊(H + 2P − K) / S⌋ + 1
H is the input side length, K the kernel size, P the padding and S the stride. A stride above 1 shrinks the output, i.e. downsamples.
그림 3

The weights of an edge-detection kernel: a positive centre surrounded by negative neighbours, highly sensitive to abrupt brightness changes (hover a cell for its weight)

-1-1-1-18-1-1-1-1

응용 분야

  • Feature-extraction backbones: the bottom of nearly every CNN, replacing hand-designed operators such as SIFT and HOG
  • Edge and gradient detection: classic kernels like Sobel and Laplacian remain in use for image processing and industrial inspection
  • Upsampling and generation: transposed convolution serves segmentation, super-resolution and decoders
  • Large-receptive-field modelling: dilated convolution is used in semantic segmentation, audio and sequence modelling

흔한 오해

  • Convolution is not synonymous with “blurring”. Blur is merely one kernel among many; the same sliding mechanism with different weights sharpens, detects edges or responds to a specific orientation.
  • Fewer parameters does not mean less computation. Weight sharing collapses the parameter count, but every location is still evaluated, so the cost grows with image area.
  • Bigger kernels are not automatically better. Modern designs prefer stacking several 3×3 kernels to reach the same receptive field with fewer parameters and more nonlinearity.

핵심 용어

Kernel / filter
The small weight matrix that is learned
Stride
How many pixels the window jumps each step
Padding
Adding zeros at the border to control output size
Receptive field
The input region that one output pixel depends on

참고문헌