Ingestion & Chunking

Retrieval can't fix bad chunks. Ingestion and chunking decide 80% of RAG quality — before any embedding exists. This reel is the craft.

▶ Watch this reel

What you'll learn

  1. Document parsing
  2. Chunking strategies
  3. Size, overlap & metadata
  4. Changing documents

Remember this

Parsing (per-format traps)

Chunking strategies

StrategyHowWhen
Fixedevery N chars + overlapdemos only
Recursiveparagraphs → sentences → wordsdefault
Semanticcut where embedding distance peaksprecision, costs a pass
Structuralfollow headings/sectionswell-structured docs

Size & overlap

Metadata per chunk

source · section · date · author · ACL → filtering, citations, security, freshness.

Changing documents

Code: Ingestion pipeline: recursive chunks + metadata + incremental upsert

import hashlib
from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,          # tokens ~ size dial: measure on golden Qs
    chunk_overlap=80,        # stitches boundaries
    separators=["\n## ", "\n# ", "\n\n", "\n", ". ", " "],
)

def chunk_id(doc_id: str, pos: int, text: str) -> str:
    # deterministic → re-ingest of same doc REPLACES its chunks
    h = hashlib.sha1(text.encode()).hexdigest()[:10]
    return f"{doc_id}:{pos}:{h}"

def ingest(doc_id: str, text: str, meta: dict):
    seen = set()
    for pos, chunk in enumerate(splitter.split_text(text)):
        cid = chunk_id(doc_id, pos, chunk)
        seen.add(cid)
        vec = embed(chunk)
        store.upsert(id=cid, vector=vec, payload={
            "text": chunk,
            "doc_id": doc_id,
            "source": meta["source"],     # citation
            "section": meta.get("section"),  # structural
            "date": meta["date"],           # freshness filter
            "acl": meta["acl"],             # security filter (GA-17)
        })
    # chunks previously stored for this doc but not in `seen` → delete
    store.delete_where({"doc_id": doc_id}, not_ids=seen)

# Same code path handles create, update, AND delete:
# changed doc → new chunk hashes → old ones purged by the last line.