Observability

LLM systems fail silently — quality decays, costs creep, latency spikes. You can't fix what you can't see: traces, cost dashboards, and quality monitoring.

▶ Watch this reel

What you'll learn

  1. Tracing LLM calls
  2. Cost dashboards
  3. Quality monitoring
  4. Tooling

Remember this

Tracing

Cost dashboards

Quality monitoring

Tooling

Code: Minimal trace decorator — the whole instrumentation habit

import time, uuid, functools

def traced(span_name):
    def deco(fn):
        @functools.wraps(fn)
        async def wrapper(*args, **kw):
            span = {
                "trace_id": kw.pop("trace_id", str(uuid.uuid4())),
                "span": span_name,
                "start": time.perf_counter(),
                "model": kw.get("model"), "prompt_hash": h(kw.get("prompt")),
            }
            try:
                result = await fn(*args, **kw)
                span.update(ok=True, tokens=getattr(result, "usage", None))
                return result
            except Exception as e:
                span.update(ok=False, error=str(e)[:200])
                raise
            finally:
                span["ms"] = round((time.perf_counter() - span["start"]) * 1000)
                emitter.send(span)          # OTLP or Langfuse/etc.
        return wrapper
    return deco

# @traced("retrieval") on search · @traced("generation") on complete —
# three decorators and your pipeline is observable end to end.