Advanced RAG
When flat chunks stop answering — graphs, agents, multi-hop, and security. The frontier toolkit for questions that span documents and permissions.
▶ Watch this reelWhat you'll learn
- GraphRAG
- Agentic RAG
- Multi-hop questions
- Security & caching
Remember this
- GraphRAG: entity/relation extraction at ingest; graph traversal for cross-document relationship questions, vectors for local ones
- Agentic RAG = retrieval as a tool under model control (judgment for latency/cost); multi-hop = fixed decompose-retrieve-synthesize for compound questions
- Ship-ready = ACL filters inside retrieval (never post-filter) + semantic cache with high similarity threshold and tenant-scoped keys
GraphRAG
- Ingest: LLM-extract entities + relations → knowledge graph beside vector index.
- Local questions → vectors; global/relationship questions → graph traversal / community summaries.
- Start flat; add graph when logs show relationship-class questions.
Agentic RAG
- Retrieval becomes a TOOL the agent invokes: decide → search → read → re-search/answer/clarify.
- Wins: judgment, no wasted retrievals, adaptive depth.
- Costs: latency, non-determinism, harder eval. Fixed decomposition (multi-hop) is the middle ground.
Multi-hop
- Decompose → retrieve per sub-question → synthesize.
- Evaluate EACH hop; answer must cite every hop's sources (citation count ≥ hop count).
Security & caching
- ACL: chunk-level metadata, filtered in the retrieval query by asking user. Never post-filter.
- Semantic cache: q-embedding → answer, threshold ≥ ~0.97, key includes tenant.
- Stage 4 complete: ingest → chunk → embed → retrieve → generate → ship.
Code: Advanced RAG guardrails: ACL filter + semantic cache + decomposition
from dataclasses import dataclass
# --- security: filter INSIDE retrieval --------------------------------
def retrieve(q, user) -> list[Chunk]:
return vdb.search(
embed(q), k=20,
where={"acl": {"$in": user.permission_groups}}, # GA-09 filters
)
# NEVER: fetch all → generate → strip disallowed from the answer.
# The LLM has already seen what you tried to hide.
# --- semantic cache: high threshold, tenant-scoped key ----------------
@dataclass
class CacheEntry:
q_vec: list[float]
answer: str
tenant: str
CACHE: list[CacheEntry] = []
def ask_cached(question: str, user) -> str:
qv = embed(question)
for e in CACHE: # small: brute force is fine
if e.tenant == user.tenant and cosine(qv, e.q_vec) >= 0.97:
return e.answer # near-duplicate only!
answer = generate(history, question, retrieve(question, user))
CACHE.append(CacheEntry(qv, answer, user.tenant))
return answer
# --- multi-hop: fixed decomposition ------------------------------------
def ask_multi_hop(compound: str) -> str:
subs = llm(f"Split into standalone sub-questions: {compound}")
results = [retrieve_and_answer(s) for s in subs] # eval each hop!
return llm(f"Synthesize with citations to ALL sources: {results}")