Performance & Cost
The same answer can cost 10× more or arrive 5× slower depending on engineering. Caching, routing, streaming, fallbacks — the optimization stack, honestly prioritized.
▶ Watch this reelWhat you'll learn
- Caching layers
- Model routing
- Latency & resilience
- Quotas & throughput
Remember this
- Three caches compound: static-first prompts for provider cache reuse, exact response cache, high-threshold semantic cache — with versioned keys and TTLs
- Model routing sends easy traffic to cheap lanes with EXTERNAL verification and policy exceptions; PTU covers the steady baseline, on-demand the peaks
- Latency = streaming + parallelism + small contexts; resilience = multi-provider fallback with constant response shapes and idempotent retries; capacity pressure sheds load by design
Caching
- Prompt cache: static-first assembly (positional).
- Response cache: exact match, tenant-scoped keys.
- Semantic cache: high threshold (≥0.95), TTLs, versioned keys, never cache errors/abstentions.
Model routing
- Cheap lane first, EXTERNAL verification, escalate on miss.
- Policy exceptions: safety/side-effects/compliance stay flagship.
- Quality monitored per route (GA-24).
Latency & resilience
- Streaming everywhere · gather() independent calls · smallest sufficient context.
- Multi-provider fallback, constant response shapes, idempotent retries.
Quotas & throughput
- PTU for steady baselines, on-demand for peaks; token buckets; queue-and-degrade ladder designed in advance; load-test quarterly.
Code: The request pipeline with every lever in place
async def handle(q, user):
key = cache_key(q, user.tenant, prompt_version())
if hit := response_cache.get(key): # exact
return hit
if near := semantic_cache.search(embed(q), 0.95):
return near # paraphrase
route = triage(q) # cheap | strong
try:
async with token_bucket(user): # smoothing
stream = route.model.stream(assemble_static_first(q))
return await emit_sse(stream, # first token <1s
on_complete=lambda r: (
response_cache.set(key, r),
semantic_cache.add(q, r)))
except ProviderUnavailable:
return await fallback_provider.handle(q) # same shape, degraded note