Retrieval Quality

Retrieval is a recall problem wearing a precision costume. Five techniques — top-k, reranking, rewriting, HyDE, hybrid — that decide whether the right chunk survives.

▶ Watch this reel

What you'll learn

  1. Top-k & the recall ceiling
  2. Reranking
  3. Query rewriting & HyDE
  4. Hybrid retrieval

Remember this

Top-k

Reranking (two-stage)

Query rewriting

Hybrid retrieval

Code: Retrieval stack: hybrid fetch → rerank → cut

from ranx import compare  # golden-set eval

def retrieve(question: str, k: int = 5) -> list[Chunk]:
    # 1 · rewrite for precision (cheap LLM call)
    q = llm(f"Rewrite as a precise standalone search query: {question}")

    # 2 · HYBRID fetch — vector + BM25, fused by rank
    vec_hits = vdb.search(embed(q), k=50)
    bm25_hits = bm25.search(q, k=50)
    fused = reciprocal_rank_fusion([vec_hits, bm25_hits])[:50]

    # 3 · RERANK with a cross-encoder
    pairs = [[question, c.text] for c in fused]   # original q, not rewrite
    scores = cross_encoder.predict(pairs)
    ranked = [c for _, c in sorted(zip(scores, fused), reverse=True)]

    return ranked[:k]   # only these enter the prompt

# --- golden-set loop: measure before you tune --------------------
# for each (question, answer_chunk) in golden_set:
#     assert answer_chunk in retrieve(question, k=20)
# recall@20 too low? → hybrid on, rewrite on, raise fetch width
# recall fine but answers wrong? → that's GA-16's stage, not this one