Vector DB Fundamentals

Exact search on a million vectors would take minutes. ANN finds the answer in milliseconds — here's the honest trade.

▶ Watch this reel

What you'll learn

  1. What & why
  2. Why not a plain SQL table
  3. ANN algorithms
  4. Recall vs speed
  5. Metadata filtering
  6. Hybrid search

Remember this

What & why

Why not a plain SQL table

ANN algorithms

- M = links/node, ef = search width → recall vs memory/insert speed.

Recall vs speed

Metadata filtering

Hybrid search

Code: Hybrid search with Qdrant-style semantics + RRF

import numpy as np

# --- setup (any ANN store: Qdrant/Chroma/pgvector all share this shape)
# collection.upsert(ids, vectors=vecs, payloads=[{text, tenant, year}...])

def vector_search(query_vec, k=50):
    return ann_index.query(query_vec, top_k=k)      # your ANN leg

def keyword_search(query, k=50):
    return bm25_index.search(query, top_k=k)        # your keyword leg

def rrf_fuse(rankings, k=60):
    """Reciprocal Rank Fusion — robust to different score scales."""
    scores = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            scores[doc_id] = scores.get(doc_id, 0) + 1.0 / (k + rank + 1)
    return sorted(scores, key=scores.get, reverse=True)

# --- one hybrid query ---------------------------------------------
vec_hits = vector_search(embed(query), k=50)
key_hits = keyword_search(query, k=50)
final = rrf_fuse([vec_hits, key_hits])[:10]

# --- recall check against brute force (your golden queries) --------
def recall_at5(ann_hits, exact_hits):
    return len(set(ann_hits[:5]) & set(exact_hits[:5])) / 5

# exact: sims = corpus_u @ query_u; top = argsort(-sims)[:k]
# assert recall_at5(ann, exact) >= 0.99  → tune M/ef until this passes