RAG vs Fine-tune vs Prompt
Three levers change model behavior: what you SAY, what it KNOWS, and what it IS. Choosing wrong wastes weeks — this reel is the decision framework architects actually use.
▶ Watch this reelWhat you'll learn
- The decision framework
- Fine-tuning, SFT, LoRA
- Distillation & small models
- Local & open models
Remember this
- Three levers: prompt (behavior, free), RAG (facts, instant updates), fine-tune (weights, permanent) — decide in that order; default stack is prompt + RAG
- SFT teaches with curated examples (quality over quantity); LoRA/PEFT trains tiny adapters — 0.1% of weights, one GPU, hot-swappable per request
- Distillation moves a teacher's capability into a cheap student; pair it with routing and the frontier model only sees the traffic that pays for it
- Open weights + Ollama (dev) / vLLM (prod, paged attention) give you an OpenAI-compatible socket you own — swap backends via base_url alone
The framework
1. Knowledge current/private/changing? → RAG (GA-09/GA-10). 2. Behavior still failing with good prompt + facts? → fine-tune. 3. Otherwise → prompt harder — the underused lever. Golden rule: fine-tune style/format, RAG facts. Complements.
Fine-tuning mechanics
- SFT on hundreds–thousands of CURATED pairs; data curation = 80% of the project.
- LoRA/PEFT: freeze base, train tiny adapters (r=16 typical). MBs, one GPU, hot-swap per request.
- Risks: catastrophic forgetting, you now own versions/deployment (GA-26), needs eval harness (GA-20).
Distillation
Teacher outputs → student training data. Gate student on golden set before traffic. Pair with routing: cheap classifier sends ~15% to frontier.
Local & open
- Ollama: dev/prototype, OpenAI-compatible on localhost:11434.
- vLLM: production serving, paged attention + continuous batching → throughput per GPU.
- When local: data-boundary rules, predictable volume capex, residency.
- When not: frontier-reasoning gap — measure on YOUR evals; the gap closes quarterly.
Cross-links
GA-03/04 (prompting, structured output), GA-09/10 (RAG), GA-20 (evals), GA-25/30 (cost), GA-26 (deployment), GA-29 (managed alternatives), AG-13 (context budget).
Code: The lever ladder
LEVER CHANGES COST REVERSIBLE? UPDATES
prompt behavior ~0 instant per call
RAG knowledge retrieval instant update docs
fine-tune weights $ + GPUs retrain per release
def choose(p):
if p.knowledge_is_private_or_changing: return "RAG"
if fails_despite_prompt_and_facts(p):
return "FINE-TUNE (LoRA) on curated examples"
return "prompt + RAG ← 80% of production answers