行列演算と線形写像
行列の積は「掛けて足す」計算ではなく、空間そのものを一度に書き換える操作だ
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
定義
A matrix is a rectangular table of numbers, and it represents a linear map: an input vector passes through the map to produce an output vector, and that map is exactly "matrix multiplication". Multiplying matrices composes two maps back to back, and the shape rule (m×n)(n×p)=(m×p) says when this is legal. Eigenvalues and the singular value decomposition break a complicated map into elementary moves: rotate, stretch, rotate.
直観的な理解
Lay a sheet of graph paper printed on a rubber membrane on the table, grab it and stretch, rotate and slant it: the grid lines deform, yet they always stay straight, parallel lines stay parallel, and the origin does not move. This kind of "no bending" deformation is a linear map, and a matrix is the operating manual for that deforming machine: read off its columns and each one says where an axis arrow ends up.
One and the same linear map can be read two ways: as a geometric deformation, or as a matrix multiplication
The covariance matrix of some Gaussian data: the diagonal holds each feature’s variance and the off-diagonal the pairwise correlations. Strongly correlated features are highly redundant in dimensionality reduction (hover for values)
仕組み
- 01
Shape first: ask whether the product is even legal
The only hard requirement is that the left matrix’s columns equal the right matrix’s rows: (m×n)·(n×p) → (m×p). A large share of deep-learning errors are simply "the shapes do not line up here". Broadcasting silently replicates data along size-1 dimensions — it makes code run, but it can also make results quietly wrong.
- 02
Multiplication is composition: apply one map, then another
C = A·B means "first apply B, then apply A". Because order matters, matrix multiplication is not commutative: A·B and B·A generally differ, just as rotating-then-translating is not the same as translating-then-rotating.
- 03
Eigenvalues and SVD: find the skeleton of the map
An eigenvector keeps its direction under the map and is merely scaled; the scale factor is its eigenvalue, defined only for square matrices. The singular value decomposition goes further for any matrix: it factors the map into "rotate · stretch along axes · rotate", and the stretches are the singular values. Keeping the leading terms yields the best low-rank approximation — the mathematical basis for dimensionality reduction, compression and recommendation.
- 04
Why GPUs are built for matrix multiplication
An n×n matrix multiplication performs roughly 2n³ floating-point operations, and every output element is independent of the others — the ideal load for parallel hardware. A GPU runs the same multiply-add across thousands of simple cores at once, turning minutes of work into milliseconds. Tensor cores hard-wire this further into silicon, so "how fast matrix multiplication runs" largely sets the cost of training and inference.
Principal component analysis of an image dataset: the leading directions already explain most of the "energy", which is why low-rank approximation works
応用場面
- A neural-network layer: multiplying the input vector by a weight matrix is one linear map
- Dimensionality reduction and compression: SVD a data matrix and keep only the leading singular values to approximate it
- Recommenders: low-rank factorisation of the user–item rating matrix fills in the missing entries
- Computer graphics: translation, rotation, scaling and projection are all matrices; composing them is just chained multiplication
よくある誤解
- Matrix multiplication is not commutative. A·B is generally not B·A; swapping them means an entirely different map, so the order must be watched in both code and derivations.
- Eigenvalues are defined only for square matrices, yet most real matrices are not square. To reason about a map of arbitrary shape you need singular values, not eigenvalues.
- Broadcasting lets mismatched shapes through, but it only fills in size-1 dimensions — not necessarily the one you intended. Silent wrong answers are therefore more dangerous than a crash.
重要用語
- Transpose
- Flip a matrix across its diagonal so rows become columns
- Eigenvector / eigenvalue
- A vector whose direction is unchanged by the map, and the factor by which it is scaled
- Singular value decomposition
- Writing any matrix as the product "rotate · stretch · rotate"
- Rank
- The number of independent directions the map actually spans; at most rows or columns