Eval Fundamentals
You can't improve what you don't measure. In GenAI, the model's mood is not a metric. Evals are how quality becomes engineering instead of vibes.
▶ Watch this reelWhat you'll learn
- Why evals
- The golden dataset
- Offline vs online
- Human review loops
Remember this
- Change is constant (prompts, silent model updates, living data) — evals are the only regression defense in GenAI
- Golden dataset: 50–200 expert-reviewed real questions, versioned like code, including unanswerable cases; quality over size
- Offline gates releases, online monitors drift; human review calibrates machines — measure the reviewers too
Why evals
- Change is constant: prompt iterations, silent model updates, living data.
- No eval → quality discovered via complaints.
- Rule: nothing ships without an eval run (same discipline as tests).
Golden dataset
- 50–200 REAL questions + expert reference answers; quality > size.
- Cover: common · edge · adversarial · unanswerable (abstention).
- Versioned (golden-v1, v2…); never leak into prompts/training.
- Owner + quarterly review; stale gold is worse than none.
Offline vs online
- Offline: golden set, reproducible, gates releases.
- Online: shadow judging, user feedback, canary analysis; monitors drift.
- Alert on DELTAS, not single scores.
Human review loops
- Rubric + calibration (shared cases, Cohen's kappa).
- Sample: random + flagged (low scores, thumbs-down, abstentions).
- Humans calibrate machines; review volume shrinks as judges improve.
Code: Versioned golden set + offline eval runner
# golden/v2.jsonl — versioned like code, reviewed by SMEs
{"q": "What is the refund window?", "must_contain": ["30 days"], "answerable": true}
{"q": "Who is the CEO's dentist?", "answerable": false} # abstain expected
# eval_runner.py
import json, datetime
def run_golden(pipeline, version: str) -> dict:
cases = [json.loads(l) for l in open(f"golden/{version}.jsonl")]
rows = []
for c in cases:
out = pipeline.ask(c["q"])
rows.append({
"q": c["q"],
"hit": all(k in out.answer for k in c.get("must_contain", [])),
"abstained": out.abstained,
"abstain_correct": (not c["answerable"]) == out.abstained,
})
return {
"dataset": version,
"run_at": datetime.datetime.utcnow().isoformat(),
"hit_rate": sum(r["hit"] for r in rows) / len(rows),
"abstention_accuracy": sum(r["abstain_correct"] for r in rows) / len(rows),
"rows": rows,
}
# Store every run: runs/2026-10-03T05-12.json
# Compare: current run vs previous runs on the SAME dataset version —
# a regression is a delta, not a vibe.