Retrieval Quality
Retrieval is a recall problem wearing a precision costume. Five techniques — top-k, reranking, rewriting, HyDE, hybrid — that decide whether the right chunk survives.
▶ Watch this reelWhat you'll learn
- Top-k & the recall ceiling
- Reranking
- Query rewriting & HyDE
- Hybrid retrieval
Remember this
- Tune top-k on golden questions (usually 8–20); fetch wide, narrow later — recall at the top, precision at the bottom
- Rerank: bi-encoder recalls broadly, cross-encoder scores (question, chunk) pairs jointly, keep the top few
- Vague queries → rewrite / multi-query / HyDE; exact tokens → hybrid BM25 + vector with reciprocal-rank fusion
Top-k
- k too small → recall dies (answer at k+1); too large → precision drowns.
- Golden questions plot recall vs k → typically 8–20.
- Funnel: fetch wide (20–50) → narrow later.
Reranking (two-stage)
- Bi-encoder: fast, independent embeddings → broad candidates.
- Cross-encoder: (q, chunk) jointly → one score; accurate, slow.
- Retrieve-then-rerank is the standard production upgrade.
Query rewriting
- Rewrite: vague → precise standalone query (conversation-aware).
- Multi-query: 3–5 paraphrases, union results.
- HyDE: generate hypothetical answer → embed it → vocabulary matches real chunks.
Hybrid retrieval
- Vector: paraphrase-proof, weak on exact tokens (codes, names, IDs).
- BM25: exact-token precision, blind to paraphrase.
- RRF fuses by rank (1, ½, ⅓…): scale-free, no calibration.
- Order: hybrid fetch → rerank → cut.
Code: Retrieval stack: hybrid fetch → rerank → cut
from ranx import compare # golden-set eval
def retrieve(question: str, k: int = 5) -> list[Chunk]:
# 1 · rewrite for precision (cheap LLM call)
q = llm(f"Rewrite as a precise standalone search query: {question}")
# 2 · HYBRID fetch — vector + BM25, fused by rank
vec_hits = vdb.search(embed(q), k=50)
bm25_hits = bm25.search(q, k=50)
fused = reciprocal_rank_fusion([vec_hits, bm25_hits])[:50]
# 3 · RERANK with a cross-encoder
pairs = [[question, c.text] for c in fused] # original q, not rewrite
scores = cross_encoder.predict(pairs)
ranked = [c for _, c in sorted(zip(scores, fused), reverse=True)]
return ranked[:k] # only these enter the prompt
# --- golden-set loop: measure before you tune --------------------
# for each (question, answer_chunk) in golden_set:
# assert answer_chunk in retrieve(question, k=20)
# recall@20 too low? → hybrid on, rewrite on, raise fetch width
# recall fine but answers wrong? → that's GA-16's stage, not this one