Classificação de imagens
Decidir a que categoria pertence a imagem inteira
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
Como é feita tecnicamente
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
Produtos representativos
4GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
Organizações relacionadas
Usos típicos
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
Como avaliar se funciona bem
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
Limites e dificuldades
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes
Conceitos por trás
Classificação de imagens
Não diga “gatos têm bigodes”: mostre gatos suficientes e ele descobre sozinho
Redes neurais convolucionais
Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional
Aprendizagem supervisionada
Pares de pergunta e resposta ensinam o modelo a responder sozinho