الرؤية ذاتية الإشراف والتعلّم التقابلي متعدد الوسائط
بلا وسوم: يتعلّم الرؤية باستنتاج أي الصور متطابقة
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Self-supervised vision pretrains representations without human labels, constructing supervision from the structure of the data itself; contrastive learning pulls different augmented views of the same image together and pushes different images apart. Multimodal contrastive learning (CLIP) goes further, aligning paired images and text in one embedding space so a model can classify zero-shot using natural language.
الحدس المباشر
Classic supervised learning needs someone to label “this is a cat” image by image. Self-supervision flips the idea: crop, flip and recolour the same image twice, tell the model “these two are the same,” then show it a pile of other images and say “those are not.” To answer correctly the model is forced to ignore irrelevant variation in colour and angle and capture the object itself — all without a single label. CLIP carries the same game to image-text pairs, drawing “a photo of a dog” and its caption together in space.
The contrastive training loop: one image yields two augmented views, the contrastive loss then draws positives together and pushes negatives apart — no human labels at any point
Linear-probe / zero-shot accuracy of self-supervised versus supervised pre-training (ResNet-50 backbone; CLIP is ViT-L/14 zero-shot)
طريقة العمل
- 01
Manufacturing supervision from the data itself
Two randomly augmented views of one image form a positive pair, while the other images in the batch serve as negatives. Since the labels come from the data, they scale without bound and are not limited by human annotation speed.
- 02
The contrastive loss: InfoNCE
The loss raises the similarity of positive pairs and lowers it against negatives, with a temperature parameter controlling the sharpness. SimCLR showed all three ingredients matter: a large batch, strong augmentation and a projection head.
- 03
MoCo: scaling negatives with a queue and a momentum encoder
Instead of relying on huge batches, MoCo keeps a queue of negatives and uses a momentum-updated encoder to keep the queued features consistent, so the number of negatives can far exceed the batch size.
- 04
Aligning vision and language: CLIP
CLIP runs bidirectional contrastive learning over 400 million image-text pairs, aligning an image encoder with a text encoder. To classify, write the class names into a prompt template such as “a photo of a {}” and compare image-text similarity — no fine-tuning required, i.e. zero-shot classification.
CLIP’s two-tower structure: separate image and text encoders map into one shared embedding space, aligned by cosine similarity
CLIP’s zero-shot ability grows with model scale: ImageNet zero-shot top-1 accuracy across backbones
- Zero-shot ImageNet
مجالات الاستخدام
- Label-scarce domains: medical, industrial and remote-sensing settings pre-train self-supervised then fine-tune on few samples
- Image-text retrieval and reverse image search: CLIP-style shared embedding spaces
- Zero-shot and open-vocabulary classification: defining new classes in natural language
- Conditioning generative models: diffusion models use a CLIP text encoder to interpret prompts
مفاهيم خاطئة شائعة
- “No labels” does not mean “no data”. CLIP used 400 million image-text pairs and SimCLR needs large batches and long training — the cost merely shifts from annotation to compute.
- Self-supervised features are not universally better than supervised ones. On several dense-prediction tasks such as detection and segmentation, supervised pre-training can still be stronger or more efficient.
- Contrastive learning leans heavily on the augmentation policy. Too weak and no invariance is learned; too strong and semantics are destroyed; the hyperparameters are sensitive.
- Zero-shot ability is sensitive to prompt wording and is limited on out-of-distribution classes and fine-grained distinctions.
مصطلحات أساسية
- Contrastive learning
- Learning representations by pulling positives together and pushing negatives apart
- InfoNCE
- The standard contrastive loss; essentially a multi-class cross-entropy
- Projection head
- The MLP the contrastive loss is applied to, usually discarded after training
- Zero-shot classification
- Classifying directly with text prompts, without fine-tuning