MCP Security
MCP servers hold credentials and take actions — that's power an attacker wants. Four threat classes every agent architect must design against: poisoning, over-permission, token leaks, and rogue servers.
▶ Watch this reelWhat you'll learn
- Tool poisoning
- Over-permissioned servers
- Token & credential handling
- Defense in depth
Remember this
- Tool poisoning hides instructions in descriptions and resources — anything in context is code; surface tool lists to users and confirm side effects
- Least privilege per server AND per tool bounds every blast radius; never let secrets be tool arguments; redact at log boundaries
- Defense in depth: vet → scope → confirm → monitor → respond; no layer is expensive, and each defeats a different failure
Tool poisoning
- Instructions hidden in tool metadata / resource content / post-approval updates.
- Defenses: surface tools to users, sanitize resource content, confirm side effects, least privilege.
Over-permissioned servers
- Least privilege per server AND per tool; split write tools onto separate credentials.
- Audit and narrow grants — permissions drift up by default.
Token handling
- Never plaintext tokens in configs/dotfiles; env refs / keychain.
- Never secrets as tool arguments — resolve inside the server.
- Redact at log boundaries; TLS; verify OAuth issuers (RFC 9207); short-lived tokens; rotation playbooks.
Defense in depth
Vet → scope → confirm → monitor → respond. Keep current with official MCP security advisories.
Code: A side-effect gate: confirmation before consequence
SIDE_EFFECTS = {"send_email", "delete_file", "update_record", "make_payment"}
async def guarded_call(server, tool_name, args, user):
call = audit.log(user, server, tool_name, args) # always log
if tool_name in SIDE_EFFECTS:
# Show WHAT, not just which tool:
summary = await describe_action(server, tool_name, args)
approved = await ui.confirm(
f"{user.name}: allow '{tool_name}'?\n\n{summary}\n\n"
f"Args: {redact_secrets(args)}") # scrub before display
if not approved:
audit.log_denied(call)
return ToolError("denied by user — do NOT retry; "
"ask the user what they'd like instead")
return await server.call(tool_name, args)
# Notes:
# - redact_secrets() runs on EVERY display/log path — defense at the sink
# - denial messages tell the model not to retry (no permission-loops)
# - the audit record exists whether approved or not