Neusortierung von Ergebnissen
Kandidaten nach echter Relevanz neu ordnen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a query and a short list of recalled candidates, and returns a reordered ranking with relevance scores. Its division of labour differs from embedding retrieval: retrieval uses cheap vector neighbours to pull dozens from tens of thousands, whereas reranking applies a costlier model to score those dozens carefully, so it only runs over the candidate set.
Wie sie technisch umgesetzt wird
The typical approach is a cross-encoder: query and candidate are concatenated into one sequence and a single model outputs a relevance score; seeing both at once is more accurate than two-tower encoders, but it cannot be pre-computed and must run per pair. To control latency, a common pattern truncates with a small model first and reranks with a larger one — or simply prompts a large model to score.
Repräsentative Produkte
4Pinecone
2019Verwaltete Vektordatenbank für Ähnlichkeitssuche
Together API
2022Eine Inferenz-API für offene Modelle
Replicate
2019Community-Modelle über eine API ausführen
Hugging Face Hub
2016Die Sammelstelle für offene Modelle und Datensätze
Beteiligte Organisationen
Typische Verwendungen
- Improving evidence quality after RAG retrieval
- Final ranking in e-commerce and content search
- Selecting the most relevant passages for a QA system
- Deduplicating and choosing among candidate answers
Wie sie bewertet wird
- nDCG@k
- Discounted gain in the top-k after reranking, rewarding better placement
- MAP
- Mean of average precision over all relevant items
- MRR
- Quality of the position of the first relevant result
Grenzen und schwierige Punkte
- Cross-encoders score each candidate separately, so latency climbs steeply with candidate count
- Judgements are unstable on out-of-distribution queries, and truncating long documents causes errors
- It can only reorder what was recalled; anything retrieval missed cannot be recovered
Konzepte dahinter
Retrieval-Augmented Generation
Wissen nicht in die Parameter stopfen, sondern draußen halten und bei Bedarf nachschlagen – wie eine Open-Book-Prüfung
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Modellbewertung und Kreuzvalidierung
Genauigkeit ist die täuschendste Metrik – falsch ausgewertet, bricht alles Weitere zusammen