RAG & LLM Metrics

Faithfulness, context precision, LLM-as-judge — the metric stack that tells you which stage of your RAG is lying and when a 'better' change actually isn't.

▶ Watch this reel

What you'll learn

  1. The metric stack
  2. LLM-as-judge
  3. Regression testing
  4. Eval tooling

Remember this

Metric stack (maps to stages)

LLM-as-judge

Regression testing

Tooling

Code: Faithfulness via atomic claim checking

JUDGE = """You are a strict fact-checker.
Given CONTEXT and ANSWER:
1. Decompose ANSWER into atomic claims (short, single-fact).
2. For each claim, classify:
   SUPPORTED   — context states it
   PARTIAL     — context implies but doesn't state it
   UNSUPPORTED — context contradicts or lacks it
3. Reply as JSON: {"claims": [{"text": ..., "verdict": ...}],
                  "faithfulness": supported/total}
Grade ONLY from the CONTEXT. Unknown = UNSUPPORTED."""

def faithfulness(context: str, answer: str) -> float:
    verdicts = llm(JUDGE, context=context, answer=answer)
    claims = verdicts["claims"]
    supported = sum(1 for c in claims if c["verdict"] == "SUPPORTED")
    return supported / max(len(claims), 1)

# Threshold in CI:  faithfulness(answer) >= 0.9
# Calibrate monthly: run 50 human-graded cases through the judge,
# compute agreement — if kappa sags, the JUDGE drifted, not the app.