Chuyển đến nội dung
Bản đồ AI

Nhúng từ (Word Embeddings)

Biến từ thành tọa độ: từ đồng nghĩa tự tụ lại và ý nghĩa lần đầu có thể cộng trừ

04 Xử lý ngôn ngữ tự nhiên và mô hình ngôn ngữ lớnCơ bảnMục từ thứ 2 trong lĩnh vực

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

Word embeddings map discrete words to continuous vectors so that semantically similar words end up close together in space. Early methods (Word2Vec, GloVe) learn one fixed vector per word; contextualised methods (ELMo, BERT) instead produce a different vector for the same word depending on its context, separating its senses.

Trực giác

A one-hot encoding hands every word an enormous card with a single 1 on it: any two words are orthogonal and equally distant, so the model cannot tell that "cat" is closer to "dog" than to "guitar". Embeddings instead place each word at a coordinate on a map, so near-synonyms crowd into the same neighbourhood and even analogies become vector arithmetic: king − man + woman ≈ queen.

Hình 1

A Skip-gram training sample: the centre word is pulled towards its in-window context words and away from unrelated ones

Hình 2

The geometry of embeddings: each country-to-capital pair shares nearly the same offset vector, and the parallelogram is visible even in 2D (hover for coordinates)

Cách hoạt động

  1. 01

    From one-hot to dense vectors

    With a million-word vocabulary, one-hot is a million-dimensional vector with a single non-zero entry — wasteful in memory and blind to similarity. Embeddings compress it into a few hundred dense dimensions, retrieved by a learnable lookup matrix.

  2. 02

    Learn vectors by predicting context (Word2Vec)

    CBOW predicts the centre word from its neighbours; Skip-gram predicts the neighbours from the centre word. The guiding assumption is that words sharing contexts share meaning. A full softmax over the vocabulary is far too slow, so negative sampling reframes the task as telling real context words apart from a few random negatives, cutting the cost dramatically.

  3. 03

    Use global co-occurrence (GloVe)

    GloVe counts how often word pairs co-occur across the whole corpus, then fits the inner product of two word vectors to the logarithm of that count. It swaps Word2Vec’s local-window prediction for global matrix factorisation, which trains more stably and parallelises more easily.

  4. 04

    Contextualised embeddings (ELMo / BERT)

    A static embedding gives "bank" one vector forever, unable to separate a riverbank from a financial institution. ELMo (bidirectional LSTM) and BERT (bidirectional Transformer) generate a vector per token conditioned on the whole sentence, so the same word takes on different representations in different contexts, largely resolving polysemy.

Hình 4

Reported accuracy on word-analogy tasks: from static embeddings to subword methods, accuracy rises in magnitude (datasets and setups differ across papers; treat as orders of magnitude)

Ứng dụng

  • Semantic search and RAG: encode documents and queries into one space and take nearest neighbours
  • Recommendation and ads: inner products of user and item vectors; higher scores mean higher click likelihood
  • Clustering and topic exploration: project a corpus to 2D to reveal natural grouping of words and documents
  • Downstream input: the first-layer word representation for larger models and the substrate of every NLP task

Hiểu lầm thường gặp

  • Static embeddings cannot separate senses. The same word is forced to share one vector across its company and fruit meanings, averaging the semantics and introducing downstream errors.
  • Analogy arithmetic is not a universal law. It works only where the corpus makes the relation highly regular; king − man + woman succeeds while rarer relations often fail, and whatever gender or cultural bias the corpus carries is preserved as-is.
  • More dimensions is not automatically better. Higher dimensionality adds parameters and overfitting risk, and gains typically flatten well past a few hundred dimensions.

Thuật ngữ chính

One-hot
A sparse vector with a single 1; distinct words are fully orthogonal
Skip-gram
A training objective that predicts surrounding words from the centre
Negative sampling
Replacing full-vocabulary softmax with a few random negatives
Cosine similarity
The alignment of two vector directions, from −1 to 1

Đọc thêm