Advanced RAG

When flat chunks stop answering — graphs, agents, multi-hop, and security. The frontier toolkit for questions that span documents and permissions.

▶ Watch this reel

What you'll learn

  1. GraphRAG
  2. Agentic RAG
  3. Multi-hop questions
  4. Security & caching

Remember this

GraphRAG

Agentic RAG

Multi-hop

Security & caching

Code: Advanced RAG guardrails: ACL filter + semantic cache + decomposition

from dataclasses import dataclass

# --- security: filter INSIDE retrieval --------------------------------
def retrieve(q, user) -> list[Chunk]:
    return vdb.search(
        embed(q), k=20,
        where={"acl": {"$in": user.permission_groups}},  # GA-09 filters
    )
    # NEVER: fetch all → generate → strip disallowed from the answer.
    # The LLM has already seen what you tried to hide.

# --- semantic cache: high threshold, tenant-scoped key ----------------
@dataclass
class CacheEntry:
    q_vec: list[float]
    answer: str
    tenant: str

CACHE: list[CacheEntry] = []

def ask_cached(question: str, user) -> str:
    qv = embed(question)
    for e in CACHE:                       # small: brute force is fine
        if e.tenant == user.tenant and cosine(qv, e.q_vec) >= 0.97:
            return e.answer               # near-duplicate only!
    answer = generate(history, question, retrieve(question, user))
    CACHE.append(CacheEntry(qv, answer, user.tenant))
    return answer

# --- multi-hop: fixed decomposition ------------------------------------
def ask_multi_hop(compound: str) -> str:
    subs = llm(f"Split into standalone sub-questions: {compound}")
    results = [retrieve_and_answer(s) for s in subs]   # eval each hop!
    return llm(f"Synthesize with citations to ALL sources: {results}")