Перейти к содержимому
Атлас ИИ
01 Основы математики и статистикиНачальныйСтатья 1 в этой области

Векторы и векторные пространства

ИИ превращает всё — слова, изображения, звуки — в список чисел

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ОПРЕДЕЛЕНИЕ

A vector is an ordered list of numbers, such as (0.21, −0.83, 0.47). A vector space is the set these vectors live in, where any two vectors can be added and any vector can be scaled by a number while staying inside the set. Nearly every input, output, parameter and intermediate result in machine learning is a point in some vector space.

Интуиция

Shrink "describe a person" to three dimensions: height, weight, age. Three people become three points standing in the same three-dimensional space. A real model may use a thousand or ten thousand dimensions, yet the geometric intuition is identical: the closer two points are, the more alike the things they represent. The single most important move in modern AI is learning to translate "meaning" into "position".

Рис. 1

From raw object to comparable coordinates: the encoder is the only door into this conversion

Рис. 2

The classic structure of word embeddings: king, queen, man and woman form a parallelogram in 2D projection (hover for coordinates)

Как это работает

  1. 01

    Encode: turn an object into coordinates

    Define a set of measurable features — or let a network learn them — then assign a value to each. A sentence, an image, an audio clip is first compressed into a fixed-length array of numbers. This step caps everything that follows.

  2. 02

    Compare: measure distance and angle

    Euclidean distance measures how far apart two points are; cosine similarity measures whether they point the same way. Text retrieval relies almost entirely on cosine, because the length of a sentence should not change its meaning.

  3. 03

    Combine: addition and linear maps

    Vector addition and matrix multiplication express "combining concepts" and "rewriting the whole space". A single neural-network layer is, at heart, one linear map followed by a nonlinear squashing.

Области применения

  • Semantic search: encode documents and queries into one space, then take the nearest neighbours
  • Recommenders: take the inner product of user and item vectors; higher scores mean higher affinity
  • Image retrieval and deduplication: nearest neighbours in feature space are usually visually similar images
  • Clustering and visualisation: project high-dimensional vectors to 2D to see whether the data groups itself

Частые заблуждения

  • High dimensionality does not mean more information. Many dimensions are strongly correlated, so the effective degrees of freedom are usually far fewer than the nominal count — which is exactly what dimensionality reduction and PCA address.
  • Distances are only comparable within one encoding. Change the model or its version and the same coordinates point to entirely different meanings; they cannot be mixed.
  • Individual dimension directions are usually not interpretable. Apart from a rare few deliberately designed axes, "what does dimension 137 mean" typically has no human-readable answer.

Ключевые термины

Dimension
The number of entries in a vector
Norm
A function measuring a vector’s "length"; L2 is the common choice
Inner product
Element-wise product summed over entries; the numerator of cosine similarity
Embedding
The layer, or its output, that maps a discrete object into a continuous vector

Дополнительная литература