Image Segmentation
Label every pixel with the object it belongs to
WHAT THIS CAPABILITY MEANS
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
How it is done
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Representative products
4Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
SenseAvatar
2022Generates lip-synced digital-human video from a portrait and a voice track
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Organizations involved
Typical uses
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
How it is evaluated
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Limits and hard parts
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Concepts behind it
Semantic Segmentation
Colouring every pixel: not a box around the object, but a colouring book
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Image Representation
To a machine, a photo is nothing but stacked grids of numbers