Threat Model
An agent with tools is a PROGRAM an attacker can steer with prose. The agent-specific threat model: confused deputies, exfiltration paths, sandboxing, and scoped identity.
▶ Watch this reelWhat you'll learn
- Indirect injection via tools
- Confused deputy
- Exfiltration paths
- Sandboxing & identity
Remember this
- Tool outputs are an injection surface — label provenance, scan patterns, and enforce at the action layer where prose becomes powerless
- Confused deputy: agents borrow authority, never own it — propagate asker scopes per call, keep service accounts minimal, attribute everything
- Exfiltration is data-flow: govern crossings with egress allow-lists and per-tool payload policies, canary-test the exits, sandbox code by default, scope identity short-lived
Indirect injection via tools
- Tool results are instructions-in-waiting · label provenance · scan patterns · enforce at action layer.
- Per-tool injection test-suite in CI; structure data vs free text where possible.
Confused deputy
- Agents borrow authority per call · propagate asker scopes · minimal service accounts · three-part attribution (agent/run/asker).
- Authority gaps route to approval chains, not agent discretion.
Exfiltration paths
- Map reads × write channels · egress allow-lists · per-tool payload policies · DLP at boundaries · canary-secret drills.
- Scope by deployment: research agent (broad web, read-only) ≠ payments agent (one API).
Sandboxing & identity
- Sandboxes: no-net default, allow-listed reads, caps, ephemeral, no host creds.
- Identity: per-agent service accounts, short-lived task-scoped tokens, pre-computed blast radius.
Code: The governed tool-call boundary, one wrapper
async def guarded_tool_call(agent, tool, args, asker):
# 1 · borrowed authority: the asker's scopes, not the agent's
if not asker.may(tool, args):
return deny(f"{asker.name} lacks {tool.required_scope(args)}")
# 2 · egress governance
if dest := args.get("url") or args.get("recipient"):
if not egress.allow(dest, via=tool.name):
return deny(f"destination '{dest}' not allow-listed for {tool.name}")
# 3 · payload policy
if size(payload_of(args)) > tool.max_payload:
return deny(f"payload {size} exceeds {tool.name} limit — split or summarize")
# 4 · canary tripwire
if CANARY in str(args):
alert("canary_in_tool_args", tool=tool.name, agent=agent.id)
# 5 · execute + attribute
result = await tool.run(scoped_creds(agent, asker), args)
audit.log(agent=agent.id, run=agent.run_id, on_behalf_of=asker.id,
tool=tool.name, args_hash=h(args))
return result