Сегментация изображений
Пометить каждый пиксель объектом, к которому он относится
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
Как это устроено
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Примеры продуктов
4Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
SenseAvatar
2022Создаёт видео цифрового человека с синхронизацией губ из портрета и дорожки голоса
GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Связанные организации
Типичное применение
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
Как её оценивают
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Границы и трудности
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Концепции в основе
Семантическая сегментация
Раскрасить каждый пиксель: не рамка вокруг объекта, а раскраска
Свёрточные нейронные сети
Замена полных связей на «смотреть локально и переиспользовать один и тот же фильтр везде» — идея, сделавшая распознавание изображений рабочим
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел