RAG vs Fine-tune vs Prompt

Three levers change model behavior: what you SAY, what it KNOWS, and what it IS. Choosing wrong wastes weeks — this reel is the decision framework architects actually use.

▶ Watch this reel

What you'll learn

  1. The decision framework
  2. Fine-tuning, SFT, LoRA
  3. Distillation & small models
  4. Local & open models

Remember this

The framework

1. Knowledge current/private/changing? → RAG (GA-09/GA-10). 2. Behavior still failing with good prompt + facts? → fine-tune. 3. Otherwise → prompt harder — the underused lever. Golden rule: fine-tune style/format, RAG facts. Complements.

Fine-tuning mechanics

Distillation

Teacher outputs → student training data. Gate student on golden set before traffic. Pair with routing: cheap classifier sends ~15% to frontier.

Local & open

Cross-links

GA-03/04 (prompting, structured output), GA-09/10 (RAG), GA-20 (evals), GA-25/30 (cost), GA-26 (deployment), GA-29 (managed alternatives), AG-13 (context budget).

Code: The lever ladder

LEVER          CHANGES       COST         REVERSIBLE?   UPDATES
prompt         behavior      ~0           instant       per call
RAG            knowledge     retrieval    instant       update docs
fine-tune      weights       $ + GPUs     retrain       per release

def choose(p):
    if p.knowledge_is_private_or_changing: return "RAG"
    if fails_despite_prompt_and_facts(p):
        return "FINE-TUNE (LoRA) on curated examples"
    return "prompt + RAG  ← 80% of production answers