Bildklassifikation
Entscheiden, zu welcher Klasse das ganze Bild gehört
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
Wie sie technisch umgesetzt wird
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
Repräsentative Produkte
4GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Beteiligte Organisationen
Typische Verwendungen
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
Wie sie bewertet wird
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
Grenzen und schwierige Punkte
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes
Konzepte dahinter
Bildklassifikation
Sag ihm nicht „Katzen haben Schnurrhaare“ – zeig ihm genug Katzen, er erkennt es selbst
Faltungsnetze
Vollverbindungen durch „lokal schauen, überall denselben Filter wiederverwenden“ ersetzen – die Idee, die Bilderkennung erst brauchbar machte
Überwachtes Lernen
Paare aus Frage und Antwort lehren das Modell, selbst zu antworten