합성곱 신경망
‘국소적으로 보고, 같은 자를 어디서나 재사용’으로 완전 연결을 대체해 이미지 인식을 처음으로 실용화했다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
A convolutional neural network is designed for grid-like data such as images. It slides a set of small filters across the input, performing the same weighted sum at every position to extract local features, and composes simple features (edges, corners) into complex ones (textures, parts, objects) layer by layer. Because one filter is reused across the whole image, the parameter count is independent of image size and translation equivariance comes for free.
직관적 이해
To find a pattern in a large jigsaw, you do not need to view it all at once; you slide a small magnifier cell by cell, noting "right angle" or "arc", then use a bigger magnifier to assemble those marks into an "eye" or a "wheel". The key insight: the same pattern is recognised by the same ruler whether it appears top-left or bottom-right — that is weight sharing, and the fundamental reason a convolution uses fewer parameters than a dense layer.
A typical convolutional stack: starting from raw pixels, convolution and pooling alternate, the receptive field grows layer by layer, and the features are finally aggregated into a class decision
ImageNet top-5 error (lower is better): from 16.4% down to the 2% range, the real ten-year progression of convolutional architectures, with human performance around 5.1% for reference
작동 원리
- 01
Local receptive field: each output sees only a patch
Each output connects to a local window of the input rather than all pixels. Adjacent pixels in an image are highly correlated, so local connections suffice to capture edges and corners while cutting the number of connections sharply.
- 02
Weight sharing: one filter slides over the whole image
The same filter is reused across the entire image, so the parameter count is independent of image size. This both shrinks the model and yields translation equivariance — a cat shifted a few pixels is still a cat.
- 03
Pooling: shrink size, enlarge the receptive field
Pooling such as max pooling takes the strongest response in a local window, shrinking spatial size while preserving salient activations. It enlarges the effective receptive field of later layers and adds robustness to small shifts.
- 04
Stacking: from edges to objects
Stacking convolution layers grows the receptive field layer by layer — shallow layers see edges, deep layers see a whole face. A final dense layer or global pooling aggregates the features into a class decision.
The evolution of convolutional architectures: each generation answers the same question — how to make the network deeper and use the spatial structure of images more effectively
응용 분야
- Image classification: from handwritten digits to the millions of images in ImageNet
- Backbones for object detection and semantic segmentation, supplying hierarchical features
- Defect detection in medical imaging, remote sensing and industrial inspection
- One-dimensional convolutions are also used for spectrograms and sensor signals
흔한 오해
- A convolution’s translation invariance is approximate. Large rotations or rescalings still break it, usually addressed with data augmentation or specialised architectures.
- Depth is not a goal in itself: before residual connections, simply going deeper raised the training error (the degradation problem), which is not overfitting.
- A small parameter count does not mean a small compute cost. Convolutions are expensive at high spatial resolution in early layers, which is exactly where inference optimisation focuses.
핵심 용어
- Kernel / filter
- A set of learnable weights that slides over the input
- Receptive field
- The region of the original input that a given output covers
- Weight sharing
- Reusing one set of weights across all spatial positions
- Equivariance
- When the input shifts, the output shifts accordingly rather than changing