Aller au contenu
Atlas de l'IA

Questions-réponses et RAG

Récupère d’abord les preuves, puis répond à partir d’elles

Langage et connaissancesDébutant #07
entréeTexteTexte

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

CE QUE DÉSIGNE CETTE CAPACITÉ

Takes a question and returns an answer. It has two forms: closed-book, answering from parameters alone, and open-book, retrieving relevant passages from an external corpus first. Retrieval-augmented generation belongs to the latter, keeping knowledge in an updatable external store and usually attaching citations. It differs from free generation by having an explicit question and a verifiable target answer.

Comment c'est fait

A RAG pipeline chunks documents into vectors in an index, encodes the question the same way and takes nearest neighbours, then pastes the retrieved passages with the question into a prompt so the model can answer with citations. Refinements include hybrid dense-plus-sparse retrieval, query rewriting before retrieval, and a reranker to pick the best passages. Closed-book answering relies purely on knowledge compressed during pre-training.

Produits représentatifs

14

Perplexity

2022
Perplexity AI

Répond en cherchant, chaque réponse appuyée par des sources

Application Freemium
TexteTexte

NotebookLM

2023
Google DeepMind

Ne répond qu’à partir de vos sources, citations à l’appui

Application Freemium
TexteTableauAudioTexteAudio

ChatGPT

2022
OpenAI

La fenêtre de discussion qui a mis un grand modèle de langage entre toutes les mains

Application Freemium
TexteImageAudioTexteImageAudio

GPT-4o

2024
OpenAI

Un modèle général nativement multimodal : texte, image et audio par une même entrée

Modèle Fermé
TexteImageAudioTexteAudio

Gemini

2024
Google DeepMind

Une entrée de discussion qui réunit recherche, bureautique et modèle multimodal

Application Freemium
TexteImageAudioTexteImageAudio

Command R

2024
Cohere

Un modèle commercial conçu pour la recherche augmentée et l’usage d’outils

Modèle Poids ouverts
TexteTexte

Jamba

2024
AI21 Labs

Un modèle à poids ouverts mêlant un modèle d’espace d’états à un Transformer

Modèle Poids ouverts
TexteTexte

Kimi

2023
Moonshot AI

Un assistant de conversation chinois réputé pour les longs contextes

Modèle Fermé
TexteTexte

Hunyuan

2023
Tencent (Hunyuan)

La famille de modèles généraux de Tencent, avec des versions à poids ouverts

Modèle Poids ouverts
TexteImageTexte

Microsoft Copilot

2023
Microsoft

Une IA conversationnelle intégrée au système d’exploitation et à la bureautique

Application Freemium
TexteImageAudioTexteImage

Doubao

2023
ByteDance (Seed)

Le modèle de conversation général et l’application de ByteDance

Modèle Fermé
TexteImageTexte

LangChain

2022
LangChain

Enchaîner modèles, outils et recherche

Outil Open source

ERNIE

2019
Baidu

Un modèle chinois parti d’un préentraînement enrichi par la connaissance, version phare précoce

Modèle Fermé
TexteTexte

Pinecone

2019
Pinecone

Base vectorielle managée pour la recherche de similarité

Infrastructure Fermé

Organisations concernées

Usages typiques

  • Enterprise knowledge bases and internal Q&A
  • Web-search assistants with citations
  • Document and contract interrogation
  • Support and technical self-service

Comment on l'évalue

Exact Match
Share of answers identical to the reference
Answer F1
Token overlap with the reference, allowing partial credit
Retrieval hit rate and citation accuracy
Whether evidence was retrieved and citations actually support the claim

Limites et points difficiles

  • When the right passage is not retrieved, the model still answers confidently — hallucination in disguise
  • Multi-hop questions needing several documents break, answering only one link
  • Key evidence in the middle of a long context is the most likely to be ignored

Concepts sous-jacents