Controls
Every threat from the last reel meets its match here: redaction, guardrails, moderation, audit trails, compliance — the control plane that never jailbreaks.
▶ Watch this reelWhat you'll learn
- PII detection & redaction
- Guardrails: input & output
- Content safety services
- Audit, residency & compliance
Remember this
- Redact before you send: detect → placeholder → route → restore, with the mapping in request scope and regex+NER combined detection
- Guardrails sandwich the model — deterministic input filters and output gates, block-or-replace by content class, every trip an observable event
- Managed moderation runs as an independent pass with threshold-to-action bands; audit logs are append-only, correlated, and privacy-shaped; residency and retention are implemented, tested features
PII detection & redaction
- Detect: regex (structured) + NER (unstructured), golden-set recall.
- Replace → route → restore; mapping in request scope, never logged; redact before APIs AND logs.
Guardrails
- Sandwich: input filters (length/encoding/injection screens) → model → output gates (secret/PII screens, topic, format).
- Block-or-replace per content class; trips are observable events with counters.
- Rules data-driven, versioned, deployable without releases.
Content safety services
- Independent pass (in/out), threshold→action bands, gray-zone human review as calibration.
- Compose with own filters: standard harm classes = managed; brand/domain = yours.
Audit & compliance
- Append-only structured logs, correlation IDs, privacy-shaped (hashes + references).
- Residency by contract (regions, self-host), retention implemented AND deletion-tested.
- Map controls → OWASP entries; quarterly control matrix.
Code: The control plane as middleware, one pipeline
async def handle_request(user, text):
audit = AuditLog.start(user.id) # correlated from birth
safe_in, mapping = redact(normalize(text)) # ch.1
if trip := injection_screen(safe_in): # ch.2 tripwire
audit.guardrail_trip("injection", trip)
return refuse("I can't process that request.")
reply = await llm.complete(safe_in) # the model proposes
if secret := secret_pattern(reply): # output gates
audit.guardrail_trip("secret_leak", hash(secret))
return refuse("The response contained sensitive data.")
mod = await moderate(reply) # ch.3 independent pass
if mod.action == "block":
audit.guardrail_trip(f"moderation_{mod.category}")
return refuse("I can't provide that content.")
if mod.action == "review":
await human_review_queue.add(reply, mod, audit.id)
audit.complete(tokens=reply.usage) # ch.4 record
return restore(reply, mapping) # ch.1 round-trip