Bildsegmentierung
Jedem Pixel das zugehörige Objekt zuweisen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
Wie sie technisch umgesetzt wird
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Repräsentative Produkte
4Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
SenseAvatar
2022Erzeugt lippensynchrones Digital-Human-Video aus einem Porträt und einer Tonspur
GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Beteiligte Organisationen
Typische Verwendungen
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
Wie sie bewertet wird
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Grenzen und schwierige Punkte
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Konzepte dahinter
Semantische Segmentierung
Jedes Pixel einfärben: kein Rahmen, sondern ein Malbuch
Faltungsnetze
Vollverbindungen durch „lokal schauen, überall denselben Filter wiederverwenden“ ersetzen – die Idee, die Bilderkennung erst brauchbar machte
Digitale Bilddarstellung
Für eine Maschine ist ein Foto nichts als gestapelte Zahlenraster