Production Engineering

Agents that run for HOURS can't live in a request handler. Durable execution, queues, budgets, and the failure catalogue — the machinery of 3 a.m. reliability.

▶ Watch this reel

What you'll learn

  1. Loop limits, timeouts, retries
  2. Durable execution
  3. Queues, concurrency & locking
  4. Cost caps & failure catalogue

Remember this

Fences

Durable execution

Queues & concurrency

Cost caps & failure catalogue

Code: The fenced, durable, metered run — one skeleton

async def run_agent(job):
    state = store.load_checkpoint(job.id) or fresh_state(job)
    budget = Budget(tokens=job.token_cap, wall=job.time_cap)

    while not state.done and not budget.exhausted():
        step = plan_next(state)
        async with timeout(step.max_seconds):
            result = await idempotent(step, key=state.idem_key(step.n))
        state.apply(step, result)
        store.checkpoint(job.id, state)            # durable after every step
        budget.spend(result.tokens)

    if not state.done:
        await escalate(job, package=state.export())  # resumable exit
    return state.result

# Properties, by construction:
#   crash      → resume from last checkpoint
#   retry      → idempotent key prevents duplicates
#   spiral     → budget exits with full state
#   cost       → metered at every model call