Architect Toolkit
GenAI architecture isn't a technology problem — it's a judgment problem. This reel hands you the working architect's kit: selection matrices, NFR checklists, cost models, and the gates between POC and production.
▶ Watch this reelWhat you'll learn
- Choosing & justifying
- Architecture patterns
- NFRs & cost model
- POC to production
Remember this
- Select use-cases on value x feasibility — kill weak ones in the meeting; high-value + low-feasibility needs a human checkpoint first
- Every GenAI system composes five blocks — model, retrieval, orchestration, guardrails, observability; buy undifferentiated blocks, build the one you win on
- GenAI's NFR triangle: accuracy SLO (evals in CI), latency (TTFT + streaming), cost per completed task — optimize one, budget the other two
- POC to production runs through gates — eval, safety, scale, operate — each with an owner; demo failure modes on purpose to earn trust
Use-case selection
- Axes: value (measurable baseline, owner, annualized impact) x feasibility (verifiability, error tolerance + human loop, data readiness, latency fit).
- High value + low feasibility → don't kill, add a human checkpoint and re-score.
- Kill low-value on the record — it buys credibility for the real 'yes'.
Reference architecture
Five blocks: model endpoint · retrieval · orchestration · guardrails · observability. Build-vs-buy per block: undifferentiated → buy; standard-needs-assembly → configure; differentiating → build. Record reasons.
NFRs (GenAI-specific)
- Accuracy: SLO on a golden eval set (AG-21); evals in CI.
- Latency: TTFT + tokens/sec are user-visible (GA-07); streaming is standard; budget per interaction type.
- Cost: per completed task = LLM calls + retrieval + guardrails + infra; compare to human baseline.
- Triangle: improve one, budget the other two explicitly.
POC → production gates
G1 eval · G2 safety · G3 scale · G4 operate — each with an OWNER. Risk register: likelihood x impact → named mitigation → cites the reel that teaches it. Client communication: demo the failure modes on purpose. Confidence is the deliverable.
Cross-links
GA-01 (probabilistic nature → feasibility), GA-07 (latency anatomy), GA-09/10 (guardrails + RAG), GA-29 (buy options), AG-21/22 (evals, red-team), AG-23 (cost discipline). AG-25 reuses this whole kit for agents.
Code: The architect's one-page scorecard
USE-CASE value feas. decision owner
support-triage 11 17 POC ✓ VP Support
contract-draft 14 11 +human HITL Legal
auto-approvals 13 6 KILL (recorded)
GATES: G1 eval ✓ G2 safety ✓ G3 scale ✓ G4 operate ✓
NFRs: accuracy ≥ 92% on golden set (CI-gated)
TTFT < 800ms, streaming on
cost/task ≤ $0.08 @ 40k tasks/mo
RISKS: 5 open, 4 mitigated — each cites a reel
One page. Defensible. Signed.