本文へスキップ
AI図鑑

単語埋め込み

「単語」を「座標」に変える。類義語は自然に集まり、意味が初めて足し引きできる

04 自然言語処理と大規模言語モデル初級この領域の第 2 項目

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

Word embeddings map discrete words to continuous vectors so that semantically similar words end up close together in space. Early methods (Word2Vec, GloVe) learn one fixed vector per word; contextualised methods (ELMo, BERT) instead produce a different vector for the same word depending on its context, separating its senses.

直観的な理解

A one-hot encoding hands every word an enormous card with a single 1 on it: any two words are orthogonal and equally distant, so the model cannot tell that "cat" is closer to "dog" than to "guitar". Embeddings instead place each word at a coordinate on a map, so near-synonyms crowd into the same neighbourhood and even analogies become vector arithmetic: king − man + woman ≈ queen.

図 1

A Skip-gram training sample: the centre word is pulled towards its in-window context words and away from unrelated ones

図 2

The geometry of embeddings: each country-to-capital pair shares nearly the same offset vector, and the parallelogram is visible even in 2D (hover for coordinates)

仕組み

  1. 01

    From one-hot to dense vectors

    With a million-word vocabulary, one-hot is a million-dimensional vector with a single non-zero entry — wasteful in memory and blind to similarity. Embeddings compress it into a few hundred dense dimensions, retrieved by a learnable lookup matrix.

  2. 02

    Learn vectors by predicting context (Word2Vec)

    CBOW predicts the centre word from its neighbours; Skip-gram predicts the neighbours from the centre word. The guiding assumption is that words sharing contexts share meaning. A full softmax over the vocabulary is far too slow, so negative sampling reframes the task as telling real context words apart from a few random negatives, cutting the cost dramatically.

  3. 03

    Use global co-occurrence (GloVe)

    GloVe counts how often word pairs co-occur across the whole corpus, then fits the inner product of two word vectors to the logarithm of that count. It swaps Word2Vec’s local-window prediction for global matrix factorisation, which trains more stably and parallelises more easily.

  4. 04

    Contextualised embeddings (ELMo / BERT)

    A static embedding gives "bank" one vector forever, unable to separate a riverbank from a financial institution. ELMo (bidirectional LSTM) and BERT (bidirectional Transformer) generate a vector per token conditioned on the whole sentence, so the same word takes on different representations in different contexts, largely resolving polysemy.

図 4

Reported accuracy on word-analogy tasks: from static embeddings to subword methods, accuracy rises in magnitude (datasets and setups differ across papers; treat as orders of magnitude)

応用場面

  • Semantic search and RAG: encode documents and queries into one space and take nearest neighbours
  • Recommendation and ads: inner products of user and item vectors; higher scores mean higher click likelihood
  • Clustering and topic exploration: project a corpus to 2D to reveal natural grouping of words and documents
  • Downstream input: the first-layer word representation for larger models and the substrate of every NLP task

よくある誤解

  • Static embeddings cannot separate senses. The same word is forced to share one vector across its company and fruit meanings, averaging the semantics and introducing downstream errors.
  • Analogy arithmetic is not a universal law. It works only where the corpus makes the relation highly regular; king − man + woman succeeds while rarer relations often fail, and whatever gender or cultural bias the corpus carries is preserved as-is.
  • More dimensions is not automatically better. Higher dimensionality adds parameters and overfitting risk, and gains typically flatten well past a few hundred dimensions.

重要用語

One-hot
A sparse vector with a single 1; distinct words are fully orthogonal
Skip-gram
A training objective that predicts surrounding words from the centre
Negative sampling
Replacing full-vocabulary softmax with a few random negatives
Cosine similarity
The alignment of two vector directions, from −1 to 1

参考文献