Testing & Quality
LLM apps are nondeterministic — your TESTS must not be. pytest, deterministic mocking, structured logs, and automation that refuses bad commits.
▶ Watch this reelWhat you'll learn
- pytest fundamentals
- Mocking LLMs & APIs
- Logging done right
- Automation: ruff, black, pre-commit
Remember this
- pytest: plain functions + fixtures (shared, teardown-safe setup) + parametrize (one function, N cases); test contracts, not exact LLM phrasing
- Mock at the boundary: FakeProvider records prompts so you assert on inputs (assembly, citations); real-model quality belongs to evals, not unit tests
- Structured JSON logs with correlation IDs make production debuggable; ruff + black + pre-commit refuse bad code before review
pytest
- Plain functions + assert; fixtures injected by name; yield → teardown.
@pytest.mark.parametrize— data-driven cases, independent reports.- LLM apps: test CONTRACTS (grounded, abstains, validates), not phrasing.
Mocking LLMs
- Never call real APIs in unit tests: slow, costly, nondeterministic.
- FakeProvider pattern: deterministic replies + records prompts → assert on INPUTS (context assembly, citation rules).
- Model quality → evals (GA-20/21), not unit tests.
Logging
- Levels: DEBUG/INFO/ERROR; structured JSON; correlation/request IDs.
- Log retrieval telemetry (rewrites, scores, sources) for quality debugging.
- Never log secrets/PII; redact at the boundary.
Automation
- ruff (lint, fast, --fix) + black (format) + pre-commit (gates on every commit).
- Shift-left: same checks locally and in CI; too cheap to bypass.
Code: The quality loop: tests + fake LLM + structured logging
import logging, json, pytest
log = logging.getLogger("rag")
class StructuredFormatter(logging.Formatter):
def format(self, record):
return json.dumps({
"ts": self.formatTime(record),
"level": record.levelname,
"event": record.msg if isinstance(record.msg, str) else None,
**getattr(record, "fields", {}), # extra context rides along
})
async def ask(rag, question, request_id):
log.info("ask.start", extra={"fields": {
"request_id": request_id, "q": question[:80]}})
answer = await rag.ask(question)
log.info("ask.done", extra={"fields": {
"request_id": request_id, "abstained": answer.abstained}})
return answer
# In tests: caplog captures everything — assert on LOGGED events:
def test_abstention_is_logged(rag, caplog):
with caplog.at_level(logging.INFO, logger="rag"):
asyncio.run(ask(rag, "totally unknown thing?", "req-1"))
abstentions = [r for r in caplog.records
if getattr(r, "abstained", None) is True]
assert abstentions, "expected an abstention event in the logs"