이미지 분류
이미지 전체가 어느 부류인지 판정한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
기술적으로 구현하는 방법
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
대표 제품
4GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
관련 기관
대표적 용도
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
성능을 평가하는 방법
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
경계와 난점
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes