Segmentação de imagens
Rotular cada pixel com o objeto a que pertence
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image and outputs a mask the size of the input, labelling each pixel with its class or its instance. It is finer than detection because it gives outlines rather than boxes. Semantic segmentation only distinguishes classes (all of them are people), while instance segmentation also separates individuals (A and B are different people); the two are often lumped together as segmentation.
Como é feita tecnicamente
The classic design is an encoder–decoder: the encoder downsamples to extract semantics, the decoder upsamples to restore resolution, and skip connections carry high-resolution detail from shallow layers, with U-Net and fully convolutional networks as landmarks. Instance segmentation often detects first and predicts a mask inside each box (the Mask R-CNN line), or uses a promptable segment-anything model cued by clicks or boxes to isolate arbitrary objects.
Produtos representativos
4Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
SenseAvatar
2022Gera vídeo de humano digital com sincronia labial a partir de um retrato e uma faixa de voz
GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Organizações relacionadas
Usos típicos
- Organ and lesion delineation in medical imaging
- Land-cover classification in remote sensing
- Drivable area and obstacles for driving
- Selection masks for image editing
Como avaliar se funciona bem
- IoU / mIoU
- Intersection over union of masks, averaged over classes
- Dice coefficient
- Common in medical imaging, more sensitive than IoU for small targets
- Boundary F-score
- Judges only contour fit rather than large correct interiors
Limites e dificuldades
- Thin structures and boundaries — hair, wires, vessel tips — break or lose their thinness
- Adjacent or overlapping instances of the same class merge into one blob
- Masks are inaccurate on transparent, reflective or low-contrast materials
Conceitos por trás
Segmentação semântica
Colorir cada pixel: não uma caixa, mas um livro de colorir
Redes neurais convolucionais
Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional
Representação digital de imagens
Para uma máquina, uma foto não passa de grades de números sobrepostas