Eval Fundamentals

You can't improve what you don't measure. In GenAI, the model's mood is not a metric. Evals are how quality becomes engineering instead of vibes.

▶ Watch this reel

What you'll learn

  1. Why evals
  2. The golden dataset
  3. Offline vs online
  4. Human review loops

Remember this

Why evals

Golden dataset

Offline vs online

Human review loops

Code: Versioned golden set + offline eval runner

# golden/v2.jsonl  — versioned like code, reviewed by SMEs
{"q": "What is the refund window?", "must_contain": ["30 days"], "answerable": true}
{"q": "Who is the CEO's dentist?",    "answerable": false}   # abstain expected

# eval_runner.py
import json, datetime

def run_golden(pipeline, version: str) -> dict:
    cases = [json.loads(l) for l in open(f"golden/{version}.jsonl")]
    rows = []
    for c in cases:
        out = pipeline.ask(c["q"])
        rows.append({
            "q": c["q"],
            "hit": all(k in out.answer for k in c.get("must_contain", [])),
            "abstained": out.abstained,
            "abstain_correct": (not c["answerable"]) == out.abstained,
        })
    return {
        "dataset": version,
        "run_at": datetime.datetime.utcnow().isoformat(),
        "hit_rate": sum(r["hit"] for r in rows) / len(rows),
        "abstention_accuracy": sum(r["abstain_correct"] for r in rows) / len(rows),
        "rows": rows,
    }

# Store every run:  runs/2026-10-03T05-12.json
# Compare: current run vs previous runs on the SAME dataset version —
# a regression is a delta, not a vibe.