RAG Architecture

RAG: the highest-ROI pattern in GenAI. One pipeline: ingest, chunk, embed, retrieve, generate. This reel is the map.

▶ Watch this reel

What you'll learn

  1. Why RAG
  2. End-to-end pipeline
  3. Naive vs advanced RAG
  4. RAG vs long context

Remember this

Why RAG

End-to-end pipeline

1. Ingest — sources → documents (formats: GA-14) 2. Chunk — documents → pieces + metadata (highest-leverage craft) 3. Embed & store — chunks → vectors → vector DB 4. Retrieve — question → top-k (hybrid/rerank: GA-15) 5. Generate — question + context → cited answer (GA-16)

Each stage separately measurable → you know WHICH stage fails.

Naive vs advanced

RAG vs long context

Code: RAG, end to end — the whole pipeline in one file

from openai import OpenAI

client = OpenAI()
COLLECTION = "handbook"

# --- ingest + chunk + embed (offline, incremental) ----------------
def index_document(doc_id: str, text: str):
    for i, chunk in enumerate(recursive_chunks(text, size=500, overlap=80)):
        vec = client.embeddings.create(
            model="text-embedding-3-small", input=chunk).data[0].embedding
        store.upsert(
            id=f"{doc_id}:{i}",
            vector=vec,
            payload={"doc_id": doc_id, "chunk": i, "text": chunk,
                     "source": doc_id, "version": doc_version(doc_id)},
        )

# --- retrieve (hybrid + filter, per GA-09/GA-15) ------------------
def retrieve(question: str, k: int = 8) -> list[str]:
    qv = client.embeddings.create(
        model="text-embedding-3-small", input=question).data[0].embedding
    hits = store.search(qv, k=k, where={"tenant": current_tenant()})
    return [h.payload["text"] for h in rerank(question, hits)[:k]]

# --- generate (grounded + cited) -----------------------------------
def ask(question: str) -> str:
    context = retrieve(question)
    prompt = f"""Answer using ONLY the context below.
If the answer isn't there, say "I don't have that information."
Cite sources as [source].

<context>
{''.join(f'<doc src="{i}">{c}</doc>' for i, c in enumerate(context))}
</context>

Question: {question}"""
    return client.chat.completions.create(
        model="gpt-4o", temperature=0,
        messages=[{"role": "user", "content": prompt}],
    ).choices[0].message.content

# That's the whole pattern. Everything after this reel improves
# one stage at a time — measure before you tune.