질의응답과 검색 증강
먼저 근거를 검색한 뒤 그것에 근거해 답한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a question and returns an answer. It has two forms: closed-book, answering from parameters alone, and open-book, retrieving relevant passages from an external corpus first. Retrieval-augmented generation belongs to the latter, keeping knowledge in an updatable external store and usually attaching citations. It differs from free generation by having an explicit question and a verifiable target answer.
기술적으로 구현하는 방법
A RAG pipeline chunks documents into vectors in an index, encodes the question the same way and takes nearest neighbours, then pastes the retrieved passages with the question into a prompt so the model can answer with citations. Refinements include hybrid dense-plus-sparse retrieval, query rewriting before retrieval, and a reranker to pick the best passages. Closed-book answering relies purely on knowledge compressed during pre-training.
대표 제품
14Perplexity
2022검색하며 답하고, 모든 답에 출처를 붙인다
NotebookLM
2023제공한 자료만 근거로 답하고 출처를 하나씩 제시한다
ChatGPT
2022대규모 언어 모델을 누구나 쓰는 대화 창으로 만들었다
GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Gemini
2024검색·오피스·멀티모달 모델을 하나의 대화 창구로 모았다
Command R
2024검색 증강과 도구 호출을 위해 설계된 상용 모델
Jamba
2024상태공간 모델과 Transformer를 혼합한 오픈웨이트 모델
Kimi
2023긴 문맥 처리에 강한 중국어 대화 어시스턴트
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
Microsoft Copilot
2023대화형 AI를 운영체제와 오피스 앱에 심었다
Doubao
2023바이트댄스의 범용 대화 모델과 앱
LangChain
2022모델·도구·검색을 연결
ERNIE
2019지식 강화 사전학습에서 출발한 중국어 모델, 초기 대표 버전
Pinecone
2019관리형 벡터 데이터베이스와 유사도 검색
관련 기관
대표적 용도
- Enterprise knowledge bases and internal Q&A
- Web-search assistants with citations
- Document and contract interrogation
- Support and technical self-service
성능을 평가하는 방법
- Exact Match
- Share of answers identical to the reference
- Answer F1
- Token overlap with the reference, allowing partial credit
- Retrieval hit rate and citation accuracy
- Whether evidence was retrieved and citations actually support the claim
경계와 난점
- When the right passage is not retrieved, the model still answers confidently — hallucination in disguise
- Multi-hop questions needing several documents break, answering only one link
- Key evidence in the middle of a long context is the most likely to be ignored