Computer Vision
Turning pixels into objects, scenes and 3D structure
OVERVIEW
- All Concepts
- 6
- Beginner
- 2
- Intermediate
- 3
- Expert
- 1
To a computer an image is just a grid of numbers; the hard part of vision has never been storing it but understanding it. This domain begins with the locality and weight sharing of convolution and connects classification, detection, segmentation, generation and 3D reconstruction, explaining what bottleneck each landmark architecture actually broke.
Questions this domain answers
- Q1
What does a convolution kernel actually learn?
- Q2
How do classification, detection and segmentation differ?
- Q3
How does a machine recover 3D from 2D images?
Concepts in this domain
- 01Image RepresentationBeginnerTo a machine, a photo is nothing but stacked grids of numbers
- 02Convolution OperationsBeginnerOne small stencil swept across the image finds edges, textures and shapes
- 03Image ClassificationIntermediateDon’t tell it “cats have whiskers” — show it enough cats and it works it out
- 04Object DetectionIntermediateFrom “what is in the image” to “what, where, and how many”
- 05Semantic SegmentationIntermediateColouring every pixel: not a box around the object, but a colouring book
- 06Self-Supervised Vision & Contrastive Multimodal LearningExpertNo labels needed: learning to see by working out which images are the same