Agent evaluation, observability and productionization
Build a quality model, evaluation set, trace, cost and capacity model, then release agents with rollout, rollback and human takeover controls.
Agent quality model
Measure more than answer accuracy. Failures can occur in retrieval, tool selection, argument generation, execution, state transitions or final presentation.
- Model: structured-output pass rate and factual consistency.
- Retrieval: Recall@K, MRR and context precision.
- Tools: tool-selection accuracy and argument validation rate.
- Task: completion rate, step success rate and retries.
- System: P50/P95 latency, throughput, errors and availability.
- Business: time saved, conversion, satisfaction and human handoff.
- Security: blocked violations, false positives and leakage incidents.
Evaluation sets and metrics
Build the set from real tasks and production Badcases, not only ideal developer-written questions. Each case should define input, expected result, allowed tools, unacceptable behavior and scoring rules.
- Facts and multi-document synthesis.
- Tool calls with invalid arguments or unavailable permissions.
- No-answer cases that must refuse instead of inventing.
- Unauthorized requests, prompt injection and malicious documents.
- Timeouts, conflicting results and downstream failures.
Combine automated checks for structure and rules, calibrated model grading for semantic quality, and human review for complex or high-risk cases.
Trace and observability
Every task needs a searchable trace connecting request, model, retrieval, tools, state and delivery.
- Request ID, user, tenant, model and version.
- Input summary, prompt version and context sources.
- Each state, tool, argument, result, duration and error.
- Tokens, retries, cost estimate and approvals.
- Final answer, citations, task status and user feedback.
Latency, cost and capacity
- Route simple tasks to smaller models and escalate only when needed.
- Summarize and retrieve only relevant context.
- Parallelize independent read-only tools with concurrency limits.
- Cache embeddings, retrieval and stable queries with permission-aware invalidation.
- Use streaming for responsiveness, but never hide final failure.
- Define fallback models, degraded modes and human handoff.
Capacity planning must include entry QPS, steps per request, model time, tool concurrency, context tokens, downstream limits and retry amplification.
Security and human takeover
- Input layer: detect malicious or untrusted content.
- Pre-execution layer: check tool allowlists and permissions.
- Execution layer: limit resources, network and side effects.
- Output layer: check sensitive data and business format.
- Require structured confirmation for deletion, payment and outbound messages.
- Freeze repeated failures or loops and preserve the complete trace.
Release and rollback
Version the model, prompt, tool schema, retrieval index and evaluation set together. Regress offline first, then run a small canary while watching quality, cost, latency and security.
Rollback must restore the associated prompt, routing, tool protocol and index version, not only application code. Side-effecting tasks also need compensation and human runbooks.
Production checklist
- Offline evaluation includes Badcases, no-answer cases and attacks.
- Every request has a trace that identifies model, retrieval, tool and state failures.
- Step, timeout, retry, rate-limit, circuit-breaker, fallback and cancellation controls exist.
- High-risk tools have authorization, approval, idempotency, audit and compensation.
- Prompt, model, schema and index versions can be rolled back.
- Quality, cost, latency, errors, success and security events have alerts.
Continue with interview questions and the practical project.