Production Engineering
Agents that run for HOURS can't live in a request handler. Durable execution, queues, budgets, and the failure catalogue — the machinery of 3 a.m. reliability.
▶ Watch this reelWhat you'll learn
- Loop limits, timeouts, retries
- Durable execution
- Queues, concurrency & locking
- Cost caps & failure catalogue
Remember this
- Fences everywhere: iteration/wall/call budgets composed per level, timeouts at every boundary, transient-only idempotent retries — budget exits carry resumable state
- Durable execution checkpoints after every step; replay-safe idempotent effects; waiting becomes free, multi-day tasks become ordinary
- Queues smooth and prioritize; version-stamped single-writer artifacts prevent silent corruption; cost meters run per-run, per-tenant, and system-wide with pre-decided degradation
Fences
- Budgets: iterations/wall/calls, nested per level · exits carry resumable state.
- Timeouts at every boundary, composed.
- Retries: transient-only, backoff+jitter, idempotent actions.
Durable execution
- Checkpoint after every step (step+result+state-delta) · replay-safe effects.
- Temporal / Durable Functions / framework checkpointing.
- Waiting is free → multi-day tasks, human gates as suspends.
Queues & concurrency
- Jobs: enqueue → priority queue → workers with provider semaphores.
- Dead-letter after N attempts; visibility timeouts > checkpoint interval.
- Artifacts: single-writer + version stamps + advisory locks.
Cost caps & failure catalogue
- Per-run caps, per-tenant quotas, system circuit breakers; 80% alerts, pre-decided degradation ladder.
- Catalogue: signature → cause → fix → prevention for every incident.
Code: The fenced, durable, metered run — one skeleton
async def run_agent(job):
state = store.load_checkpoint(job.id) or fresh_state(job)
budget = Budget(tokens=job.token_cap, wall=job.time_cap)
while not state.done and not budget.exhausted():
step = plan_next(state)
async with timeout(step.max_seconds):
result = await idempotent(step, key=state.idem_key(step.n))
state.apply(step, result)
store.checkpoint(job.id, state) # durable after every step
budget.spend(result.tokens)
if not state.done:
await escalate(job, package=state.export()) # resumable exit
return state.result
# Properties, by construction:
# crash → resume from last checkpoint
# retry → idempotent key prevents duplicates
# spiral → budget exits with full state
# cost → metered at every model call