Skip to content
AI Atlas

Question Answering & RAG

Retrieve the evidence first, then answer from it

Language & knowledgeBeginner #07
inTextText

WHAT THIS CAPABILITY MEANS

Takes a question and returns an answer. It has two forms: closed-book, answering from parameters alone, and open-book, retrieving relevant passages from an external corpus first. Retrieval-augmented generation belongs to the latter, keeping knowledge in an updatable external store and usually attaching citations. It differs from free generation by having an explicit question and a verifiable target answer.

How it is done

A RAG pipeline chunks documents into vectors in an index, encodes the question the same way and takes nearest neighbours, then pastes the retrieved passages with the question into a prompt so the model can answer with citations. Refinements include hybrid dense-plus-sparse retrieval, query rewriting before retrieval, and a reranker to pick the best passages. Closed-book answering relies purely on knowledge compressed during pre-training.

Representative products

14

Perplexity

2022
Perplexity AI

Answers built as it searches, each one backed by sources

App Freemium
TextText

NotebookLM

2023
Google DeepMind

Answers only from the sources you give it, with citations

App Freemium
TextTableAudioTextAudio

ChatGPT

2022
OpenAI

The chat window that put a large language model in everyone’s hands

App Freemium
TextImageAudioTextImageAudio

GPT-4o

2024
OpenAI

A natively multimodal general model, with text, image and audio through one door

Model Closed
TextImageAudioTextAudio

Gemini

2024
Google DeepMind

One chat entry point that gathers search, office apps and a multimodal model

App Freemium
TextImageAudioTextImageAudio

Command R

2024
Cohere

A commercial model built for retrieval augmentation and tool use

Model Open weights
TextText

Jamba

2024
AI21 Labs

An open-weight model that mixes a state-space model with a Transformer

Model Open weights
TextText

Kimi

2023
Moonshot AI

A Chinese chat assistant known for long-context handling

Model Closed
TextText

Hunyuan

2023
Tencent (Hunyuan)

Tencent’s general model family, with open-weight versions

Model Open weights
TextImageText

Microsoft Copilot

2023
Microsoft

Conversational AI woven into the operating system and office apps

App Freemium
TextImageAudioTextImage

Doubao

2023
ByteDance (Seed)

ByteDance’s general chat model and application

Model Closed
TextImageText

LangChain

2022
LangChain

Chain models, tools and retrieval together

Tool Open source

ERNIE

2019
Baidu

A Chinese model that began with knowledge-enhanced pretraining, an early landmark version

Model Closed
TextText

Pinecone

2019
Pinecone

A managed vector database for similarity search

Infrastructure Closed

Organizations involved

Typical uses

  • Enterprise knowledge bases and internal Q&A
  • Web-search assistants with citations
  • Document and contract interrogation
  • Support and technical self-service

How it is evaluated

Exact Match
Share of answers identical to the reference
Answer F1
Token overlap with the reference, allowing partial credit
Retrieval hit rate and citation accuracy
Whether evidence was retrieved and citations actually support the claim

Limits and hard parts

  • When the right passage is not retrieved, the model still answers confidently — hallucination in disguise
  • Multi-hop questions needing several documents break, answering only one link
  • Key evidence in the middle of a long context is the most likely to be ignored

Concepts behind it