Controls

Every threat from the last reel meets its match here: redaction, guardrails, moderation, audit trails, compliance — the control plane that never jailbreaks.

▶ Watch this reel

What you'll learn

  1. PII detection & redaction
  2. Guardrails: input & output
  3. Content safety services
  4. Audit, residency & compliance

Remember this

PII detection & redaction

Guardrails

Content safety services

Audit & compliance

Code: The control plane as middleware, one pipeline

async def handle_request(user, text):
    audit = AuditLog.start(user.id)               # correlated from birth

    safe_in, mapping = redact(normalize(text))    # ch.1
    if trip := injection_screen(safe_in):         # ch.2 tripwire
        audit.guardrail_trip("injection", trip)
        return refuse("I can't process that request.")

    reply = await llm.complete(safe_in)           # the model proposes

    if secret := secret_pattern(reply):           # output gates
        audit.guardrail_trip("secret_leak", hash(secret))
        return refuse("The response contained sensitive data.")
    mod = await moderate(reply)                   # ch.3 independent pass
    if mod.action == "block":
        audit.guardrail_trip(f"moderation_{mod.category}")
        return refuse("I can't provide that content.")
    if mod.action == "review":
        await human_review_queue.add(reply, mod, audit.id)

    audit.complete(tokens=reply.usage)            # ch.4 record
    return restore(reply, mapping)                # ch.1 round-trip