본문으로 건너뛰기
AI 도감

텍스트 임베딩과 의미 검색

문장을 벡터로 바꿔 의미가 가까운 것을 찾는다

언어와 지식중급 #09
입력텍스트표

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Encodes a piece of text into a fixed-length vector so that semantically similar texts sit close together, then retrieves related items by nearest-neighbour lookup. Output is a ranked result list (item ids and scores), not prose. Unlike question answering it only finds possibly relevant material; it does not compose the answer.

기술적으로 구현하는 방법

The mainstream design is a two-tower model: queries and documents are compressed by encoders, and training pushes true pairs above random negatives, typically with contrastive learning and large batches. At query time all document vectors are pre-computed into an index and an approximate nearest-neighbour search returns the top-k in milliseconds. Sparse and dense representations are often mixed to combine keyword hits with semantic recall.

대표 제품

6

관련 기관

대표적 용도

  • Search over enterprise documents and code
  • The recall stage of RAG pipelines
  • Deduplication, clustering and topic discovery
  • Recommendation and similar-content entry points

성능을 평가하는 방법

Recall@k
Whether the top-k contain all relevant documents
nDCG
Discounted cumulative gain that accounts for rank position
MRR
Mean reciprocal rank of the first relevant result

경계와 난점

  • Semantic similarity is not relevance: near neighbours may merely share wording
  • Cross-domain or cross-lingual use degrades noticeably without adaptation
  • Chunking long documents severs context, and answers at chunk boundaries are easily missed

뒤에 있는 개념