画像セグメンテーション
各画素がどの物体に属するかを割り当てる
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
技術的にどう実現するか
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
代表的な製品
4Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
SenseAvatar
2022一枚の人物画像と音声からリップシンクしたデジタルヒューマン動画を生成する
GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
関連する組織
代表的な用途
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
どう評価するか
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
限界と難しさ
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials