画像分類
画像全体がどのクラスかを判定する
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
技術的にどう実現するか
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
代表的な製品
4GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
関連する組織
代表的な用途
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
どう評価するか
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
限界と難しさ
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes