본문으로 건너뛰기
AI 도감

이미지 분할

각 픽셀이 어느 물체에 속하는지 지정한다

시각 이해입문 #13
입력이미지이미지

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.

기술적으로 구현하는 방법

The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.

대표 제품

4

관련 기관

대표적 용도

  • Organ and lesion delineation in medical imaging
  • Land-cover classification in remote sensing
  • Drivable area and obstacles for driving
  • Selection masks for image editing

성능을 평가하는 방법

IoU / mIoU
Intersection over union of masks, averaged over classes
Dice coefficient
Common in medical imaging, more sensitive than IoU for small targets
Boundary F-score
Judges only contour fit rather than large correct interiors

경계와 난점

  • Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
  • Adjacent or overlapping instances of the same class merge into one blob
  • Masks are inaccurate on transparent, reflective or low-contrast materials

뒤에 있는 개념