Testing & Quality

LLM apps are nondeterministic — your TESTS must not be. pytest, deterministic mocking, structured logs, and automation that refuses bad commits.

▶ Watch this reel

What you'll learn

  1. pytest fundamentals
  2. Mocking LLMs & APIs
  3. Logging done right
  4. Automation: ruff, black, pre-commit

Remember this

pytest

Mocking LLMs

Logging

Automation

Code: The quality loop: tests + fake LLM + structured logging

import logging, json, pytest

log = logging.getLogger("rag")

class StructuredFormatter(logging.Formatter):
    def format(self, record):
        return json.dumps({
            "ts": self.formatTime(record),
            "level": record.levelname,
            "event": record.msg if isinstance(record.msg, str) else None,
            **getattr(record, "fields", {}),   # extra context rides along
        })

async def ask(rag, question, request_id):
    log.info("ask.start", extra={"fields": {
        "request_id": request_id, "q": question[:80]}})
    answer = await rag.ask(question)
    log.info("ask.done", extra={"fields": {
        "request_id": request_id, "abstained": answer.abstained}})
    return answer

# In tests: caplog captures everything — assert on LOGGED events:
def test_abstention_is_logged(rag, caplog):
    with caplog.at_level(logging.INFO, logger="rag"):
        asyncio.run(ask(rag, "totally unknown thing?", "req-1"))
    abstentions = [r for r in caplog.records
                   if getattr(r, "abstained", None) is True]
    assert abstentions, "expected an abstention event in the logs"