セマンティックセグメンテーション
すべての画素に色を塗る —— 枠を描くのではなく塗り絵のように
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
定義
Semantic segmentation assigns a class label to every pixel, producing a label map the same size as the input. It answers “what does each pixel belong to,” yielding precise outlines rather than coarse rectangles.
直観的な理解
Detection draws a rectangle saying “there’s a cat”; segmentation traces the cat’s every contour. If detection is drawing boxes on a photo, segmentation is a colouring book: every pixel must decide which colour it belongs to.
The taxonomy of segmentation: whether same-class instances are separated divides semantic, instance and panoptic segmentation
mIoU progression on PASCAL VOC 2012: from the fully convolutional FCN to dilated convolution and encoder-decoder designs
仕組み
- 01
Where the three segmentations diverge
Semantic segmentation labels only the class (all “people” share one colour); instance segmentation additionally separates individuals (person A differs from person B); panoptic segmentation unifies the two — countable objects by instance, background by class.
- 02
Fully convolutional networks and upsampling
Replacing the classifier’s fully connected tail with convolutions keeps the spatial structure; transposed convolution or interpolation then restores the original resolution. Segmentation thus becomes pixel-wise classification instead of image-level classification.
- 03
U-Net and its skip connections
The encoder downsamples to gain semantics, the decoder upsamples to restore detail, and skip connections wire the shallow, high-resolution encoder features straight into the matching decoder layers, recovering the boundary information lost to downsampling — decisive in label-scarce medical segmentation.
- 04
Mask R-CNN: detection plus a mask branch
Mask R-CNN adds a parallel pixel-wise mask branch to Faster R-CNN, emitting a binary mask per proposal and thereby achieving instance segmentation — an elegant extension of the detection framework.
重要公式
IoU = |A ∩ B| / |A ∪ B| , mIoU = mean over classesThe symmetric structure of U-Net: the encoder downsamples, the decoder upsamples, and skip connections join features of equal resolution across the two
応用場面
- Medical imaging: outlining organs and lesions
- Autonomous driving: drivable-area and lane segmentation
- Remote-sensing mapping: land-cover classification
- Image editing: matting, background replacement and portrait segmentation
よくある誤解
- Semantic segmentation does not separate instances of the same class. Two adjacent people form one “person” region; separating them requires instance segmentation.
- Pixel accuracy is dominated by background. When sky fills most of the frame, predicting sky everywhere scores highly — which is why mIoU is preferred over pixel accuracy.
- Boundaries are always the hardest part. More downsampling sharpens semantics but blurs edges, so upsampling and skip connections constantly trade semantics against detail.
重要用語
- Semantic / instance / panoptic
- Class → class + instance → the two unified
- Transposed convolution
- An upsampling operation common in segmentation decoders
- Skip connection
- Routing shallow high-resolution features into deep layers to preserve boundaries
- mIoU
- The mean of per-class IoU, the primary segmentation metric