Threat Model

An agent with tools is a PROGRAM an attacker can steer with prose. The agent-specific threat model: confused deputies, exfiltration paths, sandboxing, and scoped identity.

▶ Watch this reel

What you'll learn

  1. Indirect injection via tools
  2. Confused deputy
  3. Exfiltration paths
  4. Sandboxing & identity

Remember this

Indirect injection via tools

Confused deputy

Exfiltration paths

Sandboxing & identity

Code: The governed tool-call boundary, one wrapper

async def guarded_tool_call(agent, tool, args, asker):
    # 1 · borrowed authority: the asker's scopes, not the agent's
    if not asker.may(tool, args):
        return deny(f"{asker.name} lacks {tool.required_scope(args)}")

    # 2 · egress governance
    if dest := args.get("url") or args.get("recipient"):
        if not egress.allow(dest, via=tool.name):
            return deny(f"destination '{dest}' not allow-listed for {tool.name}")

    # 3 · payload policy
    if size(payload_of(args)) > tool.max_payload:
        return deny(f"payload {size} exceeds {tool.name} limit — split or summarize")

    # 4 · canary tripwire
    if CANARY in str(args):
        alert("canary_in_tool_args", tool=tool.name, agent=agent.id)

    # 5 · execute + attribute
    result = await tool.run(scoped_creds(agent, asker), args)
    audit.log(agent=agent.id, run=agent.run_id, on_behalf_of=asker.id,
              tool=tool.name, args_hash=h(args))
    return result