Fragen und Antworten mit Retrieval
Zuerst Belege abrufen, dann daraus antworten
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a question and returns an answer. It has two forms: closed-book, answering from parameters alone, and open-book, retrieving relevant passages from an external corpus first. Retrieval-augmented generation belongs to the latter, keeping knowledge in an updatable external store and usually attaching citations. It differs from free generation by having an explicit question and a verifiable target answer.
Wie sie technisch umgesetzt wird
A RAG pipeline chunks documents into vectors in an index, encodes the question the same way and takes nearest neighbours, then pastes the retrieved passages with the question into a prompt so the model can answer with citations. Refinements include hybrid dense-plus-sparse retrieval, query rewriting before retrieval, and a reranker to pick the best passages. Closed-book answering relies purely on knowledge compressed during pre-training.
Repräsentative Produkte
14Perplexity
2022Antwortet beim Suchen, jede Antwort mit Quellen belegt
NotebookLM
2023Antwortet nur aus den von dir gelieferten Quellen, mit Belegen
ChatGPT
2022Das Chatfenster, das ein großes Sprachmodell für alle zugänglich machte
GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2024Ein Chat-Eingang, der Suche, Office-Apps und ein multimodales Modell bündelt
Command R
2024Ein kommerzielles Modell für Retrieval-Augmentierung und Werkzeugnutzung
Jamba
2024Ein Open-Weights-Modell, das ein Zustandsraummodell mit einem Transformer mischt
Kimi
2023Ein chinesischer Chat-Assistent, bekannt für langen Kontext
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Microsoft Copilot
2023Konversations-KI, eingebettet in Betriebssystem und Office-Apps
Doubao
2023ByteDances allgemeines Dialogmodell und App
LangChain
2022Modelle, Werkzeuge und Retrieval verketten
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Pinecone
2019Verwaltete Vektordatenbank für Ähnlichkeitssuche
Beteiligte Organisationen
Typische Verwendungen
- Enterprise knowledge bases and internal Q&A
- Web-search assistants with citations
- Document and contract interrogation
- Support and technical self-service
Wie sie bewertet wird
- Exact Match
- Share of answers identical to the reference
- Answer F1
- Token overlap with the reference, allowing partial credit
- Retrieval hit rate and citation accuracy
- Whether evidence was retrieved and citations actually support the claim
Grenzen und schwierige Punkte
- When the right passage is not retrieved, the model still answers confidently — hallucination in disguise
- Multi-hop questions needing several documents break, answering only one link
- Key evidence in the middle of a long context is the most likely to be ignored
Konzepte dahinter
Retrieval-Augmented Generation
Wissen nicht in die Parameter stopfen, sondern draußen halten und bei Bedarf nachschlagen – wie eine Open-Book-Prüfung
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Prompt-Engineering und Alignment
Ein Modell hilfreich, ehrlich und harmlos zu machen ist schwieriger, als es einfach größer zu machen