Xếp hạng lại kết quả
Sắp xếp lại ứng viên theo độ liên quan thực
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a query and a short list of recalled candidates, and returns a reordered ranking with relevance scores. Its division of labour differs from embedding retrieval: retrieval uses cheap vector neighbours to pull dozens from tens of thousands, whereas reranking applies a costlier model to score those dozens carefully, so it only runs over the candidate set.
Làm ra sao về mặt kỹ thuật
The typical approach is a cross-encoder: query and candidate are concatenated into one sequence and a single model outputs a relevance score; seeing both at once is more accurate than two-tower encoders, but it cannot be pre-computed and must run per pair. To control latency, a common pattern truncates with a small model first and reranks with a larger one — or simply prompts a large model to score.
Sản phẩm tiêu biểu
4Pinecone
2019Cơ sở dữ liệu vector được quản lý cho tìm kiếm tương đồng
Together API
2022API suy luận cho các mô hình mở
Replicate
2019Chạy mô hình do cộng đồng đăng qua API
Hugging Face Hub
2016Nơi quy tụ mô hình và dữ liệu mở
Tổ chức liên quan
Cách dùng tiêu biểu
- Improving evidence quality after RAG retrieval
- Final ranking in e-commerce and content search
- Selecting the most relevant passages for a QA system
- Deduplicating and choosing among candidate answers
Đánh giá nó tốt hay không thế nào
- nDCG@k
- Discounted gain in the top-k after reranking, rewarding better placement
- MAP
- Mean of average precision over all relevant items
- MRR
- Quality of the position of the first relevant result
Ranh giới và điểm khó
- Cross-encoders score each candidate separately, so latency climbs steeply with candidate count
- Judgements are unstable on out-of-distribution queries, and truncating long documents causes errors
- It can only reorder what was recalled; anything retrieval missed cannot be recovered
Các khái niệm đằng sau
Sinh văn bản tăng cường truy xuất
Thay vì nhồi kiến thức vào tham số, hãy để nó ở ngoài và tra khi cần — như thi mở sách
Cơ chế chú ý (Attention)
Mọi vị trí đều có thể nhìn thẳng vào mọi vị trí khác và phân bổ chú ý theo mức liên quan
Đánh giá mô hình và kiểm định chéo
Độ chính xác là chỉ số dễ đánh lừa nhất — đánh giá sai thì mọi thứ sau đó đều vô nghĩa