RAG Architecture
RAG: the highest-ROI pattern in GenAI. One pipeline: ingest, chunk, embed, retrieve, generate. This reel is the map.
▶ Watch this reelWhat you'll learn
- Why RAG
- End-to-end pipeline
- Naive vs advanced RAG
- RAG vs long context
Remember this
- RAG = fresh + private + grounded answers without retraining; evidence is inspectable
- Pipeline: ingest → chunk → embed → retrieve → generate — each stage separable and measurable
- Naive RAG fails predictably (chunks, ambiguity, depth); long context complements rather than replaces retrieval
Why RAG
- LLM knowledge: frozen at training, outside your perimeter, fills gaps with invention.
- RAG: retrieve relevant facts at query time → prompt with evidence.
- Freshness (re-index = current) · Privacy (data stays home) · Grounding (answers cite real text — inspectable).
End-to-end pipeline
1. Ingest — sources → documents (formats: GA-14) 2. Chunk — documents → pieces + metadata (highest-leverage craft) 3. Embed & store — chunks → vectors → vector DB 4. Retrieve — question → top-k (hybrid/rerank: GA-15) 5. Generate — question + context → cited answer (GA-16)
Each stage separately measurable → you know WHICH stage fails.
Naive vs advanced
- Naive: fixed chunks + one vector search. Fails on: chunking, ambiguous queries, multi-doc/multi-hop answers.
- Advanced = diagnosed failure mode → named fix (better chunks, rewrite, hybrid, rerank, agentic loop).
RAG vs long context
- Stuffing wins: small bounded corpus, task needs most of it.
- RAG wins: large corpus, slice per question — cost, latency, attention quality.
- Hybrid: retrieve broadly → stuff winners into the long window.
Code: RAG, end to end — the whole pipeline in one file
from openai import OpenAI
client = OpenAI()
COLLECTION = "handbook"
# --- ingest + chunk + embed (offline, incremental) ----------------
def index_document(doc_id: str, text: str):
for i, chunk in enumerate(recursive_chunks(text, size=500, overlap=80)):
vec = client.embeddings.create(
model="text-embedding-3-small", input=chunk).data[0].embedding
store.upsert(
id=f"{doc_id}:{i}",
vector=vec,
payload={"doc_id": doc_id, "chunk": i, "text": chunk,
"source": doc_id, "version": doc_version(doc_id)},
)
# --- retrieve (hybrid + filter, per GA-09/GA-15) ------------------
def retrieve(question: str, k: int = 8) -> list[str]:
qv = client.embeddings.create(
model="text-embedding-3-small", input=question).data[0].embedding
hits = store.search(qv, k=k, where={"tenant": current_tenant()})
return [h.payload["text"] for h in rerank(question, hits)[:k]]
# --- generate (grounded + cited) -----------------------------------
def ask(question: str) -> str:
context = retrieve(question)
prompt = f"""Answer using ONLY the context below.
If the answer isn't there, say "I don't have that information."
Cite sources as [source].
<context>
{''.join(f'<doc src="{i}">{c}</doc>' for i, c in enumerate(context))}
</context>
Question: {question}"""
return client.chat.completions.create(
model="gpt-4o", temperature=0,
messages=[{"role": "user", "content": prompt}],
).choices[0].message.content
# That's the whole pattern. Everything after this reel improves
# one stage at a time — measure before you tune.