Question Answering & RAG
Retrieve the evidence first, then answer from it
WHAT THIS CAPABILITY MEANS
Takes a question and returns an answer. It has two forms: closed-book, answering from parameters alone, and open-book, retrieving relevant passages from an external corpus first. Retrieval-augmented generation belongs to the latter, keeping knowledge in an updatable external store and usually attaching citations. It differs from free generation by having an explicit question and a verifiable target answer.
How it is done
A RAG pipeline chunks documents into vectors in an index, encodes the question the same way and takes nearest neighbours, then pastes the retrieved passages with the question into a prompt so the model can answer with citations. Refinements include hybrid dense-plus-sparse retrieval, query rewriting before retrieval, and a reranker to pick the best passages. Closed-book answering relies purely on knowledge compressed during pre-training.
Representative products
14Perplexity
2022Answers built as it searches, each one backed by sources
NotebookLM
2023Answers only from the sources you give it, with citations
ChatGPT
2022The chat window that put a large language model in everyone’s hands
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
Command R
2024A commercial model built for retrieval augmentation and tool use
Jamba
2024An open-weight model that mixes a state-space model with a Transformer
Kimi
2023A Chinese chat assistant known for long-context handling
Hunyuan
2023Tencent’s general model family, with open-weight versions
Microsoft Copilot
2023Conversational AI woven into the operating system and office apps
Doubao
2023ByteDance’s general chat model and application
LangChain
2022Chain models, tools and retrieval together
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Pinecone
2019A managed vector database for similarity search
Organizations involved
Typical uses
- Enterprise knowledge bases and internal Q&A
- Web-search assistants with citations
- Document and contract interrogation
- Support and technical self-service
How it is evaluated
- Exact Match
- Share of answers identical to the reference
- Answer F1
- Token overlap with the reference, allowing partial credit
- Retrieval hit rate and citation accuracy
- Whether evidence was retrieved and citations actually support the claim
Limits and hard parts
- When the right passage is not retrieved, the model still answers confidently — hallucination in disguise
- Multi-hop questions needing several documents break, answering only one link
- Key evidence in the middle of a long context is the most likely to be ignored
Concepts behind it
Retrieval-Augmented Generation
Rather than cramming knowledge into parameters, leave it outside and look it up on demand — an open-book exam instead of a closed-book one
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger