Observability
LLM systems fail silently — quality decays, costs creep, latency spikes. You can't fix what you can't see: traces, cost dashboards, and quality monitoring.
▶ Watch this reelWhat you'll learn
- Tracing LLM calls
- Cost dashboards
- Quality monitoring
- Tooling
Remember this
- Traces are the flight recorder: correlated IDs, sampled errors fully, PII-shaped storage, span-level detail that localizes failures to retrieval or generation
- Cost dashboards slice spend per user/feature/model with delta alerts, runtime caps, and cost-quality pairing on the same time axis
- Quality monitoring closes the loop: feedback and shadow evals feed the golden set through SME triage; pick observability per ecosystem with OpenTelemetry underneath
Tracing
- One trace per request: retrieval/generation/tool spans, correlated IDs.
- 100% of errors + sampled successes; golden-set runs always traced.
- PII-shaped storage: hashes, access control, retention.
Cost dashboards
- Dollars per user/tenant · per feature · per model/prompt version.
- Delta alerts + runtime caps; cost×quality on one axis; forecasting run-rate.
Quality monitoring
- Feedback → triage (traces localize the span) → SME-confirmed golden cases → gated fixes.
- Shadow evals on sampled traffic; input-drift watch predicts quality drift.
Tooling
- Langfuse (OSS/self-host) · LangSmith (LangGraph) · Azure Monitor (MAF) — ecosystem match.
- Always emit OpenTelemetry underneath; keep the provider replaceable.
Code: Minimal trace decorator — the whole instrumentation habit
import time, uuid, functools
def traced(span_name):
def deco(fn):
@functools.wraps(fn)
async def wrapper(*args, **kw):
span = {
"trace_id": kw.pop("trace_id", str(uuid.uuid4())),
"span": span_name,
"start": time.perf_counter(),
"model": kw.get("model"), "prompt_hash": h(kw.get("prompt")),
}
try:
result = await fn(*args, **kw)
span.update(ok=True, tokens=getattr(result, "usage", None))
return result
except Exception as e:
span.update(ok=False, error=str(e)[:200])
raise
finally:
span["ms"] = round((time.perf_counter() - span["start"]) * 1000)
emitter.send(span) # OTLP or Langfuse/etc.
return wrapper
return deco
# @traced("retrieval") on search · @traced("generation") on complete —
# three decorators and your pipeline is observable end to end.