Vector DB Options
Fifteen products, one decision matrix. Pick by constraints, not by hype.
▶ Watch this reelWhat you'll learn
- pgvector
- Azure AI Search
- Qdrant / Chroma / FAISS
- Pinecone & managed SaaS
- Relational built-ins
- The selection matrix
Remember this
- pgvector = consolidation default; Azure AI Search = Microsoft-estate enterprise; Qdrant/Chroma/FAISS = self-host spectrum; Pinecone = SaaS velocity
- Relational DBs are adding vectors — verify per-version before proposing
- Benchmark your corpus, write the ADR, keep the exit open via a thin retrieval interface
pgvector
- Postgres + vectors. Default for apps/prototypes/multi-tenant to millions of vectors.
- Watch: HNSW memory, vacuum habits. Exit when measurements demand.
Azure AI Search
- Managed, Entra-integrated, hybrid + semantic ranking, AI enrichment.
- Wins: Azure estates, compliance-driven clients. Model SKU/capacity pricing.
Open-source spectrum
- FAISS: library (indexes only — you build the DB).
- Chroma: embedded, zero-ops — prototypes → small prod.
- Qdrant: self-hosted production server — filtering-friendly HNSW, replication.
Pinecone & managed SaaS
- Zero ops, serverless scaling, ongoing rent, data leaves your VPC.
- 12-month crossover math decides vs self-host.
Relational built-ins
- SQL Server / Oracle / others adding vector search — verify per exact version.
- 'Already licensed' wins procurement rooms when it clears the technical bar.
Selection matrix
- Score: scale · ops · security/data-residency · cost (steady-state math).
- Benchmark YOUR corpus (recall + filtered QPS) · write the ADR · keep a thin Retriever interface so the store is replaceable.
Code: The thin retrieval interface — swap stores without tears
from typing import Protocol, Any
class Hit(BaseModel):
id: str
score: float
payload: dict
class Retriever(Protocol):
def upsert(self, ids: list[str], vecs, payloads: list[dict]) -> None: ...
def search(self, vec, k: int = 10, where: dict | None = None) -> list[Hit]: ...
def delete(self, ids: list[str]) -> None: ...
# --- one implementation per store ---------------------------------
class QdrantRetriever:
def __init__(self, client, collection: str):
self.c, self.col = client, collection
def search(self, vec, k=10, where=None):
r = self.c.search(self.col, vec, limit=k,
query_filter=self._f(where) if where else None)
return [Hit(id=h.id, score=h.score, payload=h.payload) for h in r]
class PgvectorRetriever:
def __init__(self, engine): self.engine = engine
def search(self, vec, k=10, where=None):
sql = text("""SELECT id, 1-(embedding <=> :q) AS score, payload
FROM docs
WHERE (:tenant IS NULL OR tenant_id = :tenant)
ORDER BY embedding <=> :q LIMIT :k""")
... # same Hit shape
# --- app code depends on the Protocol, never a store --------------
def answer_question(q: str, retriever: Retriever):
hits = retriever.search(embed(q), k=5, where={"tenant": current_tenant()})
return generate(context=[h.payload["text"] for h in hits])
# Qdrant today, pgvector for a conservative client, Azure next —
# the app doesn't change.