Segmentation d’images
Étiqueter chaque pixel selon l’objet auquel il appartient
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
Comment c'est fait
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Produits représentatifs
4Gemini
2023Un modèle général nativement multimodal, conçu pour de très longs contextes
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
SenseAvatar
2022Génère une vidéo d’humain numérique synchronisée sur les lèvres à partir d’un portrait et d’une piste vocale
GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Organisations concernées
Usages typiques
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
Comment on l'évalue
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Limites et points difficiles
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Concepts sous-jacents
Segmentation sémantique
Colorier chaque pixel : non pas un cadre, mais un livre à colorier
Réseaux de neurones convolutifs
Remplacer les connexions denses par « regarder localement et réutiliser le même filtre partout » : l’idée qui a rendu la reconnaissance d’images enfin fonctionnelle
Représentation numérique de l’image
Pour une machine, une photo n’est qu’une pile de grilles de nombres