본문으로 건너뛰기
AI 도감

합성곱 신경망

‘국소적으로 보고, 같은 자를 어디서나 재사용’으로 완전 연결을 대체해 이미지 인식을 처음으로 실용화했다

03 딥러닝중급이 영역의 4번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

A convolutional neural network is designed for grid-like data such as images. It slides a set of small filters across the input, performing the same weighted sum at every position to extract local features, and composes simple features (edges, corners) into complex ones (textures, parts, objects) layer by layer. Because one filter is reused across the whole image, the parameter count is independent of image size and translation equivariance comes for free.

직관적 이해

To find a pattern in a large jigsaw, you do not need to view it all at once; you slide a small magnifier cell by cell, noting "right angle" or "arc", then use a bigger magnifier to assemble those marks into an "eye" or a "wheel". The key insight: the same pattern is recognised by the same ruler whether it appears top-left or bottom-right — that is weight sharing, and the fundamental reason a convolution uses fewer parameters than a dense layer.

그림 1

A typical convolutional stack: starting from raw pixels, convolution and pooling alternate, the receptive field grows layer by layer, and the features are finally aggregated into a class decision

Shallow layers extract edges and textures; deep layers compose them into parts and objects. Depth sets the level of abstraction.
그림 2

ImageNet top-5 error (lower is better): from 16.4% down to the 2% range, the real ten-year progression of convolutional architectures, with human performance around 5.1% for reference

작동 원리

  1. 01

    Local receptive field: each output sees only a patch

    Each output connects to a local window of the input rather than all pixels. Adjacent pixels in an image are highly correlated, so local connections suffice to capture edges and corners while cutting the number of connections sharply.

  2. 02

    Weight sharing: one filter slides over the whole image

    The same filter is reused across the entire image, so the parameter count is independent of image size. This both shrinks the model and yields translation equivariance — a cat shifted a few pixels is still a cat.

  3. 03

    Pooling: shrink size, enlarge the receptive field

    Pooling such as max pooling takes the strongest response in a local window, shrinking spatial size while preserving salient activations. It enlarges the effective receptive field of later layers and adds robustness to small shifts.

  4. 04

    Stacking: from edges to objects

    Stacking convolution layers grows the receptive field layer by layer — shallow layers see edges, deep layers see a whole face. A final dense layer or global pooling aggregates the features into a class decision.

그림 3

The evolution of convolutional architectures: each generation answers the same question — how to make the network deeper and use the spatial structure of images more effectively

1998LeNet-5The first mature example of conv + pooling + dense, for handwritten digits2012AlexNetReLU + GPU + Dropout; 16.4% ImageNet top-5 error2014VGG / GoogLeNetUniform 3×3 stacks and the Inception module; error down to 6.7%2015ResNetResidual connections make 152 layers trainable; 3.57% error2019EfficientNetCompound scaling jointly tunes depth, width and resolution

응용 분야

  • Image classification: from handwritten digits to the millions of images in ImageNet
  • Backbones for object detection and semantic segmentation, supplying hierarchical features
  • Defect detection in medical imaging, remote sensing and industrial inspection
  • One-dimensional convolutions are also used for spectrograms and sensor signals

흔한 오해

  • A convolution’s translation invariance is approximate. Large rotations or rescalings still break it, usually addressed with data augmentation or specialised architectures.
  • Depth is not a goal in itself: before residual connections, simply going deeper raised the training error (the degradation problem), which is not overfitting.
  • A small parameter count does not mean a small compute cost. Convolutions are expensive at high spatial resolution in early layers, which is exactly where inference optimisation focuses.

핵심 용어

Kernel / filter
A set of learnable weights that slides over the input
Receptive field
The region of the original input that a given output covers
Weight sharing
Reusing one set of weights across all spatial positions
Equivariance
When the input shifts, the output shifts accordingly rather than changing

참고문헌