मुख्य सामग्री पर जाएँ

शब्द एम्बेडिंग

शब्दों को निर्देशांक बनाना — समानार्थी स्वयं पास आ जाते हैं और अर्थ को जोड़ना-घटाना संभव हो जाता है

04 प्राकृतिक भाषा प्रसंस्करण और बड़े भाषा मॉडलप्रारंभिकइस क्षेत्र की प्रविष्टि 2

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

परिभाषा

Word embeddings map discrete words to continuous vectors so that semantically similar words end up close together in space. Early methods (Word2Vec, GloVe) learn one fixed vector per word; contextualised methods (ELMo, BERT) instead produce a different vector for the same word depending on its context, separating its senses.

सहज समझ

A one-hot encoding hands every word an enormous card with a single 1 on it: any two words are orthogonal and equally distant, so the model cannot tell that "cat" is closer to "dog" than to "guitar". Embeddings instead place each word at a coordinate on a map, so near-synonyms crowd into the same neighbourhood and even analogies become vector arithmetic: king − man + woman ≈ queen.

चित्र 1

A Skip-gram training sample: the centre word is pulled towards its in-window context words and away from unrelated ones

चित्र 2

The geometry of embeddings: each country-to-capital pair shares nearly the same offset vector, and the parallelogram is visible even in 2D (hover for coordinates)

कार्यप्रणाली

  1. 01

    From one-hot to dense vectors

    With a million-word vocabulary, one-hot is a million-dimensional vector with a single non-zero entry — wasteful in memory and blind to similarity. Embeddings compress it into a few hundred dense dimensions, retrieved by a learnable lookup matrix.

  2. 02

    Learn vectors by predicting context (Word2Vec)

    CBOW predicts the centre word from its neighbours; Skip-gram predicts the neighbours from the centre word. The guiding assumption is that words sharing contexts share meaning. A full softmax over the vocabulary is far too slow, so negative sampling reframes the task as telling real context words apart from a few random negatives, cutting the cost dramatically.

  3. 03

    Use global co-occurrence (GloVe)

    GloVe counts how often word pairs co-occur across the whole corpus, then fits the inner product of two word vectors to the logarithm of that count. It swaps Word2Vec’s local-window prediction for global matrix factorisation, which trains more stably and parallelises more easily.

  4. 04

    Contextualised embeddings (ELMo / BERT)

    A static embedding gives "bank" one vector forever, unable to separate a riverbank from a financial institution. ELMo (bidirectional LSTM) and BERT (bidirectional Transformer) generate a vector per token conditioned on the whole sentence, so the same word takes on different representations in different contexts, largely resolving polysemy.

चित्र 4

Reported accuracy on word-analogy tasks: from static embeddings to subword methods, accuracy rises in magnitude (datasets and setups differ across papers; treat as orders of magnitude)

उपयोग के क्षेत्र

  • Semantic search and RAG: encode documents and queries into one space and take nearest neighbours
  • Recommendation and ads: inner products of user and item vectors; higher scores mean higher click likelihood
  • Clustering and topic exploration: project a corpus to 2D to reveal natural grouping of words and documents
  • Downstream input: the first-layer word representation for larger models and the substrate of every NLP task

सामान्य भ्रांतियाँ

  • Static embeddings cannot separate senses. The same word is forced to share one vector across its company and fruit meanings, averaging the semantics and introducing downstream errors.
  • Analogy arithmetic is not a universal law. It works only where the corpus makes the relation highly regular; king − man + woman succeeds while rarer relations often fail, and whatever gender or cultural bias the corpus carries is preserved as-is.
  • More dimensions is not automatically better. Higher dimensionality adds parameters and overfitting risk, and gains typically flatten well past a few hundred dimensions.

मुख्य शब्द

One-hot
A sparse vector with a single 1; distinct words are fully orthogonal
Skip-gram
A training objective that predicts surrounding words from the centre
Negative sampling
Replacing full-vocabulary softmax with a few random negatives
Cosine similarity
The alignment of two vector directions, from −1 to 1

आगे का पठन