Agent Evals
Agent quality is more than the final answer: did it take sensible steps, call the right tools, waste nothing? Trajectory evaluation and the agent regression suite.
▶ Watch this reelWhat you'll learn
- Task success rate
- Trajectory & tool-call evaluation
- Simulated environments
- Cost, latency & regression suites
Remember this
- Task success is verified completion against rubric-graded golden tasks — including adversarial cases where success means refusing — distributed by difficulty and domain
- Trajectory evaluation decomposes HOW agents succeed: tool selection, argument accuracy, step sense, efficiency — each dimension maps to a specific fix
- Simulated environments with seeded failures make breaking things free; efficiency metrics (cost/latency/calls per task) are promotion-gated regressions
Task success rate
- Verified completion (artifact + spec + checks) · rubric-graded correctness · adversarial golden tasks (success = refuse/clarify/escalate).
- Distributions by difficulty and domain; paired with efficiency.
Trajectory evaluation
- Dimensions: tool selection, arg accuracy, step sense, efficiency vs reference path.
- Judge-calibrated (GA-21); mechanical where possible (malformed-call ratio).
- Diagnostic: each dimension maps to a fix; trajectories are autonomy evidence (GA-17).
Simulated environments
- Snapshot sandboxes · interaction mocks · adversarial seeding (failures at configurable rates).
- Incident replays become permanent regression cases.
Cost, latency & regression
- Per-task signatures: tokens, wall-clock, calls, tool calls.
- Efficiency regressions gate promotion like quality regressions.
- Suite grows at its frontier: incidents, escalations, feedback → new cases.
Code: The agent regression gate, as CI config
agent_regression:
dataset: agent_tasks/v4 # golden + adversarial + incidents
gates:
success_rate: { min: 0.90, by_difficulty: { easy: 0.98, hard: 0.60 } }
trajectory: { tool_selection: 0.85, arg_accuracy: 0.95,
step_sense: 0.80, efficiency: 0.70 }
resources: { cost_per_task_max: 1.4x_baseline,
calls_per_task_max: 1.3x_baseline }
safety: { adversarial_pass: 1.0 } # refusals where required
cadence:
pr: smoke_20_tasks
nightly: full_suite → trend_report
release: full_suite + simulation_replay
# Red gate → no promotion → last-known-good keeps serving.