Классификация изображений
Определить, к какому классу относится изображение целиком
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an image and outputs a category label or a probability over classes. It judges the whole image and gives no location, which is what separates it from detection and segmentation. The class set is fixed at training time, ranging from the thousand-odd general classes to narrow sets for a species or a defect type.
Как это устроено
Convolutional networks long dominated, and residual connections were the turning point that let them grow deep reliably. The image is resized to a fixed size, features are extracted layer by layer, and a classification head emits probabilities. Vision Transformers have since replaced convolution with patch embeddings and attention, matching it under large-scale pre-training, while self-supervised pre-training makes good results possible with little labelling.
Примеры продуктов
4GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
Связанные организации
Типичное применение
- Automatic tagging of products and photos
- Lesion screening in medical imaging
- Defect judgement in industrial inspection
- Content moderation and safety filtering
Как её оценивают
- Top-1 / Top-5 accuracy
- Share where the top or top-five guesses contain the correct class
- F1 and confusion matrix
- Reveals which two classes are being confused
- Expected calibration error
- Whether predicted confidence matches actual accuracy
Границы и трудности
- Long-tailed and rare classes are chronically missed for lack of examples
- Out-of-distribution inputs and tiny adversarial perturbations swing the prediction
- Fine-grained distinctions such as similar car models rely on texture cues and score far lower than coarse classes
Концепции в основе
Классификация изображений
Не говорите «у кошек есть усы» — покажите достаточно кошек, и он поймёт сам
Свёрточные нейронные сети
Замена полных связей на «смотреть локально и переиспользовать один и тот же фильтр везде» — идея, сделавшая распознавание изображений рабочим
Обучение с учителем
Пары «вопрос — ответ» учат модель отвечать самостоятельно