Ingestion & Chunking
Retrieval can't fix bad chunks. Ingestion and chunking decide 80% of RAG quality — before any embedding exists. This reel is the craft.
▶ Watch this reelWhat you'll learn
- Document parsing
- Chunking strategies
- Size, overlap & metadata
- Changing documents
Remember this
- Parsing is per-format: PDF reading order, OCR for scans, special handling for tables — inspect extractions before indexing
- Chunk strategies: fixed (naive) < recursive (default) < semantic/structural (precision) · sweet spot ~200–500 tokens + 50–100 overlap
- Metadata per chunk (source, section, date, ACL) turns search into an enterprise system; deterministic IDs make updates and deletes clean
Parsing (per-format traps)
- PDF: layout ≠ reading order; multi-column scramble → layout-aware parser.
- DOCX/PPTX/HTML: structured but noisy (text boxes, nav junk, boilerplate).
- Scanned: OCR only — its accuracy is your ceiling; QC before indexing.
- Tables: hardest — torn rows are unretrievable; serialize with headers or chunk whole.
Chunking strategies
| Strategy | How | When |
|---|---|---|
| Fixed | every N chars + overlap | demos only |
| Recursive | paragraphs → sentences → words | default |
| Semantic | cut where embedding distance peaks | precision, costs a pass |
| Structural | follow headings/sections | well-structured docs |
Size & overlap
- ~200–500 tokens prose sweet spot; validate with golden questions.
- Overlap 50–100 tokens: no sentence lives in only one chunk.
Metadata per chunk
source · section · date · author · ACL → filtering, citations, security, freshness.
Changing documents
- Deterministic IDs (doc:pos:hash) → upsert = clean replace.
- Diff at doc level; re-embed only changed docs.
- Deletes: cascade chunk removal — no zombie content.
Code: Ingestion pipeline: recursive chunks + metadata + incremental upsert
import hashlib
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500, # tokens ~ size dial: measure on golden Qs
chunk_overlap=80, # stitches boundaries
separators=["\n## ", "\n# ", "\n\n", "\n", ". ", " "],
)
def chunk_id(doc_id: str, pos: int, text: str) -> str:
# deterministic → re-ingest of same doc REPLACES its chunks
h = hashlib.sha1(text.encode()).hexdigest()[:10]
return f"{doc_id}:{pos}:{h}"
def ingest(doc_id: str, text: str, meta: dict):
seen = set()
for pos, chunk in enumerate(splitter.split_text(text)):
cid = chunk_id(doc_id, pos, chunk)
seen.add(cid)
vec = embed(chunk)
store.upsert(id=cid, vector=vec, payload={
"text": chunk,
"doc_id": doc_id,
"source": meta["source"], # citation
"section": meta.get("section"), # structural
"date": meta["date"], # freshness filter
"acl": meta["acl"], # security filter (GA-17)
})
# chunks previously stored for this doc but not in `seen` → delete
store.delete_where({"doc_id": doc_id}, not_ids=seen)
# Same code path handles create, update, AND delete:
# changed doc → new chunk hashes → old ones purged by the last line.