RAG & LLM Metrics
Faithfulness, context precision, LLM-as-judge — the metric stack that tells you which stage of your RAG is lying and when a 'better' change actually isn't.
▶ Watch this reelWhat you'll learn
- The metric stack
- LLM-as-judge
- Regression testing
- Eval tooling
Remember this
- Metric stack maps to stages: context recall/precision → retrieval · faithfulness → generation honesty · relevance → intent
- LLM-as-judge scales grading but carries position, verbosity, and self-preference biases — calibrate against human anchors, grade atomically
- Promote evals to CI gates (fast subset on PR, full nightly) and record model/prompt/dataset versions per run
Metric stack (maps to stages)
- Context recall/precision → retrieval stage (GA-15 fixes).
- Faithfulness → generation honesty (claim-by-claim support).
- Answer relevance → intent alignment.
- Pattern-matching triage: which metric dropped names the stage.
LLM-as-judge
- Biases: position, verbosity, self-preference.
- Mitigate: rubric-based absolute scores, randomized order, atomic claim checks, different-family judge.
- Calibrate vs human anchors (kappa) — scheduled, not once.
Regression testing
- PR gate: fast subset, thresholds on deltas.
- Nightly full golden run → dashboard + trend alerts.
- Record model version + prompt hash + dataset version per run.
Tooling
- Ragas (metrics), promptfoo (prompt testing/CI), Langfuse (tracing).
- Tools ship default judges — override and calibrate deliberately.
Code: Faithfulness via atomic claim checking
JUDGE = """You are a strict fact-checker.
Given CONTEXT and ANSWER:
1. Decompose ANSWER into atomic claims (short, single-fact).
2. For each claim, classify:
SUPPORTED — context states it
PARTIAL — context implies but doesn't state it
UNSUPPORTED — context contradicts or lacks it
3. Reply as JSON: {"claims": [{"text": ..., "verdict": ...}],
"faithfulness": supported/total}
Grade ONLY from the CONTEXT. Unknown = UNSUPPORTED."""
def faithfulness(context: str, answer: str) -> float:
verdicts = llm(JUDGE, context=context, answer=answer)
claims = verdicts["claims"]
supported = sum(1 for c in claims if c["verdict"] == "SUPPORTED")
return supported / max(len(claims), 1)
# Threshold in CI: faithfulness(answer) >= 0.9
# Calibrate monthly: run 50 human-graded cases through the judge,
# compute agreement — if kappa sags, the JUDGE drifted, not the app.