Chuyển đến nội dung
Bản đồ AI
08 Kỹ thuật, an toàn và đạo đức AICơ bảnMục từ thứ 4 trong lĩnh vực

Sinh văn bản tăng cường truy xuất

Thay vì nhồi kiến thức vào tham số, hãy để nó ở ngoài và tra khi cần — như thi mở sách

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

Retrieval-Augmented Generation (RAG) retrieves passages relevant to a question from an external knowledge base and feeds them to the model as context before it generates an answer. It moves “memory” out of the model’s parameters into an external store that can be updated at will and traced to a source, easing stale knowledge, hallucination and the inability to cite.

Trực giác

A closed-book exam demands you memorise everything, with no way to update afterwards; an open-book exam lets you bring the text and check the uncertain bits before answering. RAG gives the model a book and a librarian: the model’s reasoning answers the question, the retrieval system fetches the pages, and the book can be swapped for a fresh edition every day.

Hình 1

The RAG pipeline: from question to cited answer, retrieval, reranking and generation each play a part, and a failure anywhere drags down the final answer

Hình 2

Average performance on a BEIR-style multi-dataset retrieval benchmark (illustrative magnitude): hybrid retrieval with reranking clearly beats any single retriever

  • Recall@100
  • nDCG@10

Cách hoạt động

  1. 01

    Chunk and embed: turn documents into retrievable units

    Long documents are first split into reasonably sized chunks, then encoded by an embedding model into vectors stored in a vector database. Chunks too large add noise and dilute the signal; chunks too small sever the meaning — chunking strategy is often the first thing that decides RAG quality.

  2. 02

    Retrieve: vectors and keywords cover for each other

    Dense vector retrieval excels at semantically similar but lexically different matches, while sparse retrieval such as BM25 excels at exact keywords, IDs and proper nouns. Hybrid retrieval fuses both sets and is usually steadier than either retriever alone.

  3. 03

    Rerank: lift the truly relevant to the top

    The first pass is coarse because it must be fast. A reranker applies a slower but more accurate cross-encoder to score each question–passage pair, running only on a short candidate list, raising top-k quality a great deal at limited cost.

  4. 04

    Assemble and generate: give the answer a source

    The reranked passages plus the question go into the prompt, with an instruction to answer only from the supplied material and cite it. This reframes the model from “answering from memory” to “summarising the given evidence”, and makes every claim traceable back to a source.

Hình 3

Where fine-tuning and RAG differ: one writes knowledge into the parameters and changes behaviour, the other keeps knowledge outside and changes what the model sees

Ứng dụng

  • Enterprise knowledge Q&A: traceable answers over internal documents, tickets and manuals
  • Support and documentation assistants: putting the latest product notes into answers without retraining
  • Citation-critical settings: law, medicine and finance where every claim needs a source
  • Fast-changing knowledge: news, policy and inventory that no retraining cycle can keep up with

Hiểu lầm thường gặp

  • RAG cannot fix a model’s reasoning. It supplies facts, not logic; with all the material present, a faulty chain of reasoning still yields a wrong answer.
  • When retrieval comes up empty the model may still answer confidently. Unless the “no relevant material” case is handled explicitly, it falls back on parametric memory and fabricates just as fluently.
  • More context is not better. Stuffing in many irrelevant passages dilutes attention, raises cost and can feed noise into the answer; retrieval quality matters more than quantity.

Thuật ngữ chính

Chunking
Splitting long documents into retrievable pieces
Embedding model
A model that encodes text into vectors
Vector database
A store providing nearest-neighbour search over high-dimensional vectors
Reranking
Rescoring candidate passages with a more accurate model

Đọc thêm