Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas
El texto completo se presenta en inglés; el título y el resumen están traducidos.
DEFINICIÓN
A digital image is a multidimensional array of pixel values: a grayscale image is a 2D matrix, colour adds a channel axis (usually R, G and B), and a batch stacks them into a 4D tensor (N, C, H, W). Bit depth, colour space (sRGB / HSV / Lab) and the normalisation scheme together define what the network actually sees as input.
Intuición
A 1000×1000 colour photo is, to a computer, about 1000×1000×3 ≈ 3 million numbers. Zoom to the pixel level and the photo becomes a wall of mosaic tiles, each inscribed with three numbers between 0 and 255. What you see is “a cat”; what the machine sees is three overlaid tables of numbers describing red, green and blue intensity. The first step of computer vision is always to accept this rather than skip past it.
A grayscale image is a 2D matrix: each cell holds a 0–255 brightness value, and the bright band forms a vertical edge
The grayscale histogram of a “dark background + bright object” image: the two clear peaks correspond to the background and subject brightness levels
Cómo funciona
- 01
Sample and quantise: cutting the continuous world into a grid
Light in the world is continuous; the sensor samples colour at each grid cell and quantises intensity into integers (8 bits means 0–255). Resolution fixes how dense the grid is, bit depth how many shades each pixel can hold.
- 02
Channels and colour spaces: three coordinate systems for one pixel
RGB is device-facing, HSV decouples “colour” from “brightness”, and Lab aims at perceptual uniformity. Before colour thresholding, white balancing or style transfer, choosing the right colour space often helps more than a bigger model.
- 03
Tensor shape: feeding the image to the network
A single image is (C, H, W); a batch is (N, C, H, W). Most frameworks default to channels-first (NCHW), while many data pipelines hand over channels-last (H, W, C) — a shape mismatch is the most common beginner error.
- 04
Normalisation and augmentation: the last step before the network
Subtracting the mean and dividing by the standard deviation puts channels on a common scale and stabilises optimisation; random crops, flips and colour jitter inflate the data without changing semantics, forming the first line of defence against overfitting.
A colour image grows from 2D to 4D: two spatial axes plus a channel axis plus a batch axis, forming (N, C, H, W)
Dónde se usa
- Preprocessing pipelines: decode → resize → normalise → tensorise, the shared first step of nearly every vision model
- Medical and remote-sensing imagery: 16-bit high dynamic range and false-colour mapping, where bit depth decides whether lesions or land cover are visible at all
- Colour-centric tasks: skin detection, traffic-light recognition and white balancing rely on HSV / Lab
- Data augmentation: uniformly enlarging the training set is the cheapest route to better generalisation
Errores comunes
- Normalisation parameters must match training. Train with ImageNet statistics but infer on raw 0–255 input and the outputs quietly corrupt — no exception is raised, only accuracy falls.
- Higher resolution is not automatically better. Downsampling erases small objects, yet blind upscaling adds no information while raising compute quadratically.
- RGB is not “the colour a human sees”. sRGB is a display convention; Lab is closer to perceptually uniform. Be clear about which one a colour distance is measured in.
Términos clave
- Pixel
- The smallest sampling unit of an image, carrying one or more channel values
- Bit depth
- How many bits encode each channel; 8 bits give 256 levels
- Channel
- A distinct measurement at the same location, such as R/G/B or alpha
- Colour space
- A coordinate system for colour values, such as sRGB, HSV or Lab