Segmentación de imágenes
Etiquetar cada píxel con su objeto
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
Cómo se consigue técnicamente
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Productos representativos
4Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
SenseAvatar
2022Genera vídeo de humano digital con sincronía labial desde un retrato y una pista de voz
GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Organizaciones relacionadas
Usos típicos
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
Cómo se evalúa
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Límites y dificultades
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Conceptos detrás
Segmentación semántica
Colorear cada píxel: no un recuadro, sino un libro para colorear
Redes neuronales convolucionales
Sustituir las conexiones completas por “mirar en local y reutilizar la misma regla en todas partes”: la idea que hizo funcionar el reconocimiento de imágenes
Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas