Chuyển đến nội dung
Bản đồ AI

Phân loại hình ảnh

Xác định cả bức ảnh thuộc lớp nào

Hiểu thị giácCơ bản #11
đầu vàoHình ảnhVăn bản

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NĂNG LỰC NÀY NGHĨA LÀ GÌ

Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.

Làm ra sao về mặt kỹ thuật

Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.

Sản phẩm tiêu biểu

4

Tổ chức liên quan

Cách dùng tiêu biểu

  • Automatic tagging of products and photos
  • Lesion screening in medical imaging
  • Defect judgement in industrial inspection
  • Content moderation and safety filtering

Đánh giá nó tốt hay không thế nào

Top-1 / Top-5 accuracy
Share where the top or top-five guesses contain the correct class
F1 and confusion matrix
Reveals which two classes are being confused
Expected calibration error
Whether predicted confidence matches actual accuracy

Ranh giới và điểm khó

  • Long-tailed and rare classes are chronically missed for lack of examples
  • Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
  • Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes

Các khái niệm đằng sau