मुख्य सामग्री पर जाएँ

टेक्स्ट एम्बेडिंग और अर्थ-आधारित खोज

वाक्यों को सदिश बनाकर अर्थ के निकटतम खोजना

भाषा और ज्ञानमध्यवर्ती #09
इनपुटटेक्स्टटेबल

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Encodes a piece of text into a fixed-length vector so that semantically similar texts sit close together, then retrieves related items by nearest-neighbour lookup. Output is a ranked result list (item ids and scores), not prose. Unlike question answering it only finds possibly relevant material; it does not compose the answer.

तकनीकी रूप से कैसे

The mainstream design is a two-tower model: queries and documents are compressed by encoders, and training pushes true pairs above random negatives, typically with contrastive learning and large batches. At query time all document vectors are pre-computed into an index and an approximate nearest-neighbour search returns the top-k in milliseconds. Sparse and dense representations are often mixed to combine keyword hits with semantic recall.

प्रतिनिधि उत्पाद

6

संबंधित संस्थान

सामान्य उपयोग

  • Search over enterprise documents and code
  • The recall stage of RAG pipelines
  • Deduplication, clustering and topic discovery
  • Recommendation and similar-content entry points

इसका मूल्यांकन कैसे होता है

Recall@k
Whether the top-k contain all relevant documents
nDCG
Discounted cumulative gain that accounts for rank position
MRR
Mean reciprocal rank of the first relevant result

सीमाएँ और कठिनाइयाँ

  • Semantic similarity is not relevance: near neighbours may merely share wording
  • Cross-domain or cross-lingual use degrades noticeably without adaptation
  • Chunking long documents severs context, and answers at chunk boundaries are easily missed

इसके पीछे की अवधारणाएँ