본문으로 건너뛰기
AI 도감

시맨틱 분할

모든 픽셀에 색칠하기 — 상자를 그리는 게 아니라 색칠 공부처럼

05 컴퓨터 비전중급이 영역의 5번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Semantic segmentation assigns a class label to every pixel, producing a label map the same size as the input. It answers “what does each pixel belong to,” yielding precise outlines rather than coarse rectangles.

직관적 이해

Detection draws a rectangle saying “there’s a cat”; segmentation traces the cat’s every contour. If detection is drawing boxes on a photo, segmentation is a colouring book: every pixel must decide which colour it belongs to.

그림 1

The taxonomy of segmentation: whether same-class instances are separated divides semantic, instance and panoptic segmentation

Image segmentationImage segmentationSemantic segmentationSemantic segmentationOne class per pixelOne class per pixelSame-class instances mergedSame-class instances mergedInstance segmentationInstance segmentationSeparates individuals of the same classSeparates individuals of the…Usually built on detectionUsually built on detectionPanoptic segmentationPanoptic segmentationCountable objects by instanceCountable objects by instanceBackground regions by classBackground regions by class
그림 2

mIoU progression on PASCAL VOC 2012: from the fully convolutional FCN to dilated convolution and encoder-decoder designs

작동 원리

  1. 01

    Where the three segmentations diverge

    Semantic segmentation labels only the class (all “people” share one colour); instance segmentation additionally separates individuals (person A differs from person B); panoptic segmentation unifies the two — countable objects by instance, background by class.

  2. 02

    Fully convolutional networks and upsampling

    Replacing the classifier’s fully connected tail with convolutions keeps the spatial structure; transposed convolution or interpolation then restores the original resolution. Segmentation thus becomes pixel-wise classification instead of image-level classification.

  3. 03

    U-Net and its skip connections

    The encoder downsamples to gain semantics, the decoder upsamples to restore detail, and skip connections wire the shallow, high-resolution encoder features straight into the matching decoder layers, recovering the boundary information lost to downsampling — decisive in label-scarce medical segmentation.

  4. 04

    Mask R-CNN: detection plus a mask branch

    Mask R-CNN adds a parallel pixel-wise mask branch to Faster R-CNN, emitting a binary mask per proposal and thereby achieving instance segmentation — an elegant extension of the detection framework.

핵심 수식

IoU = |A ∩ B| / |A ∪ B| , mIoU = mean over classes
IoU measures the overlap between predicted and ground-truth regions; mIoU averages it over all classes and is the primary segmentation metric.
그림 3

The symmetric structure of U-Net: the encoder downsamples, the decoder upsamples, and skip connections join features of equal resolution across the two

Skip connections are what let U-Net beat a plain encoder-decoder on boundary detail.

응용 분야

  • Medical imaging: outlining organs and lesions
  • Autonomous driving: drivable-area and lane segmentation
  • Remote-sensing mapping: land-cover classification
  • Image editing: matting, background replacement and portrait segmentation

흔한 오해

  • Semantic segmentation does not separate instances of the same class. Two adjacent people form one “person” region; separating them requires instance segmentation.
  • Pixel accuracy is dominated by background. When sky fills most of the frame, predicting sky everywhere scores highly — which is why mIoU is preferred over pixel accuracy.
  • Boundaries are always the hardest part. More downsampling sharpens semantics but blurs edges, so upsampling and skip connections constantly trade semantics against detail.

핵심 용어

Semantic / instance / panoptic
Class → class + instance → the two unified
Transposed convolution
An upsampling operation common in segmentation decoders
Skip connection
Routing shallow high-resolution features into deep layers to preserve boundaries
mIoU
The mean of per-class IoU, the primary segmentation metric

참고문헌