本文へスキップ
AI図鑑

画像のデジタル表現

機械にとって写真とは、数字のグリッドを重ねたものにすぎない

05 コンピュータビジョン初級この領域の第 1 項目

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

A digital image is a multidimensional array of pixel values: a grayscale image is a 2D matrix, colour adds a channel axis (usually R, G and B), and a batch stacks them into a 4D tensor (N, C, H, W). Bit depth, colour space (sRGB / HSV / Lab) and the normalisation scheme together define what the network actually sees as input.

直観的な理解

A 1000×1000 colour photo is, to a computer, about 1000×1000×3 ≈ 3 million numbers. Zoom to the pixel level and the photo becomes a wall of mosaic tiles, each inscribed with three numbers between 0 and 255. What you see is “a cat”; what the machine sees is three overlaid tables of numbers describing red, green and blue intensity. The first step of computer vision is always to accept this rather than skip past it.

図 1

A grayscale image is a 2D matrix: each cell holds a 0–255 brightness value, and the bright band forms a vertical edge

12121014902101113801820512147019212116018521520519551822211951817225215218517
図 2

The grayscale histogram of a “dark background + bright object” image: the two clear peaks correspond to the background and subject brightness levels

仕組み

  1. 01

    Sample and quantise: cutting the continuous world into a grid

    Light in the world is continuous; the sensor samples colour at each grid cell and quantises intensity into integers (8 bits means 0–255). Resolution fixes how dense the grid is, bit depth how many shades each pixel can hold.

  2. 02

    Channels and colour spaces: three coordinate systems for one pixel

    RGB is device-facing, HSV decouples “colour” from “brightness”, and Lab aims at perceptual uniformity. Before colour thresholding, white balancing or style transfer, choosing the right colour space often helps more than a bigger model.

  3. 03

    Tensor shape: feeding the image to the network

    A single image is (C, H, W); a batch is (N, C, H, W). Most frameworks default to channels-first (NCHW), while many data pipelines hand over channels-last (H, W, C) — a shape mismatch is the most common beginner error.

  4. 04

    Normalisation and augmentation: the last step before the network

    Subtracting the mean and dividing by the standard deviation puts channels on a common scale and stabilises optimisation; random crops, flips and colour jitter inflate the data without changing semantics, forming the first line of defence against overfitting.

図 3

A colour image grows from 2D to 4D: two spatial axes plus a channel axis plus a batch axis, forming (N, C, H, W)

Frameworks default to NCHW; data arriving as (H, W, C) must be transposed explicitly.

応用場面

  • Preprocessing pipelines: decode → resize → normalise → tensorise, the shared first step of nearly every vision model
  • Medical and remote-sensing imagery: 16-bit high dynamic range and false-colour mapping, where bit depth decides whether lesions or land cover are visible at all
  • Colour-centric tasks: skin detection, traffic-light recognition and white balancing rely on HSV / Lab
  • Data augmentation: uniformly enlarging the training set is the cheapest route to better generalisation

よくある誤解

  • Normalisation parameters must match training. Train with ImageNet statistics but infer on raw 0–255 input and the outputs quietly corrupt — no exception is raised, only accuracy falls.
  • Higher resolution is not automatically better. Downsampling erases small objects, yet blind upscaling adds no information while raising compute quadratically.
  • RGB is not “the colour a human sees”. sRGB is a display convention; Lab is closer to perceptually uniform. Be clear about which one a colour distance is measured in.

重要用語

Pixel
The smallest sampling unit of an image, carrying one or more channel values
Bit depth
How many bits encode each channel; 8 bits give 256 levels
Channel
A distinct measurement at the same location, such as R/G/B or alpha
Colour space
A coordinate system for colour values, such as sRGB, HSV or Lab

参考文献