टेक्स्ट एम्बेडिंग और अर्थ-आधारित खोज
वाक्यों को सदिश बनाकर अर्थ के निकटतम खोजना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Encodes a piece of text into a fixed-length vector so that semantically similar texts sit close together, then retrieves related items by nearest-neighbour lookup. Output is a ranked result list (item ids and scores), not prose. Unlike question answering it only finds possibly relevant material; it does not compose the answer.
तकनीकी रूप से कैसे
The mainstream design is a two-tower model: queries and documents are compressed by encoders, and training pushes true pairs above random negatives, typically with contrastive learning and large batches. At query time all document vectors are pre-computed into an index and an approximate nearest-neighbour search returns the top-k in milliseconds. Sparse and dense representations are often mixed to combine keyword hits with semantic recall.
प्रतिनिधि उत्पाद
6Pinecone
2019समानता खोज हेतु प्रबंधित वेक्टर डेटाबेस
Hugging Face Hub
2016खुले मॉडल और डेटासेट का संगम
Transformers
2018पूर्व-प्रशिक्षित मॉडल लोड व प्रशिक्षण हेतु एक API
Together API
2022ओपन मॉडल हेतु इन्फ़रेंस API
Replicate
2019API से समुदाय के मॉडल चलाएँ
Perplexity
2022खोजते हुए उत्तर देता है, हर उत्तर के साथ स्रोत
संबंधित संस्थान
सामान्य उपयोग
- Search over enterprise documents and code
- The recall stage of RAG pipelines
- Deduplication, clustering and topic discovery
- Recommendation and similar-content entry points
इसका मूल्यांकन कैसे होता है
- Recall@k
- Whether the top-k contain all relevant documents
- nDCG
- Discounted cumulative gain that accounts for rank position
- MRR
- Mean reciprocal rank of the first relevant result
सीमाएँ और कठिनाइयाँ
- Semantic similarity is not relevance: near neighbours may merely share wording
- Cross-domain or cross-lingual use degrades noticeably without adaptation
- Chunking long documents severs context, and answers at chunk boundaries are easily missed
इसके पीछे की अवधारणाएँ
शब्द एम्बेडिंग
शब्दों को निर्देशांक बनाना — समानार्थी स्वयं पास आ जाते हैं और अर्थ को जोड़ना-घटाना संभव हो जाता है
सदिश और सदिश समष्टि
एआई हर चीज़ — शब्द, चित्र, ध्वनि — को संख्याओं की सूची बना देता है
पुनर्प्राप्ति-संवर्धित जनरेशन
ज्ञान को पैरामीटर में ठूँसने के बजाय बाहर रखकर ज़रूरत पर देखना — खुली किताब वाली परीक्षा की तरह