AI evaluation / Insight

Evaluating AI agents: measure completed work, not convincing answers

An agent can write a perfect completion message after updating the wrong record, reading forbidden data or repeating an action. Evaluate the work that actually happened, the authority used to do it and the uncertainty left behind. Answer quality is one signal; it is not the release decision.

Article

For: Engineering leads, architects and product owners preparing an AI agent for a bounded production release.

Define the unit of success before choosing a grader#

Suppose an internal assistant is asked to create a follow-up task for a particular customer. A good result is not the sentence ‘Task created’. It is one correct task, in the authorized customer record, with the approved owner and due date, backed by a retrievable receipt. The result fails if the right task appears in the wrong account, if a duplicate exists, or if the assistant read another customer’s restricted notes to produce it.

Anthropic’s guide to agent evaluations usefully distinguishes a trial’s transcript from its outcome in the environment. Apply that distinction to your own system: model output describes what the agent says happened; business-state checks establish what happened. Evaluation must cover the complete application and its tools, not just a prompt tested in isolation.

Write a task contract containing initial state, user authority, allowed operations, required final state, prohibited effects and acceptable handovers. If two competent reviewers disagree about whether a case passes, fix the contract before collecting a larger score. A fuzzy requirement does not become objective because a model assigns it a number.

Use a scorecard that cannot average away harm#

On a narrow screen, scroll the table sideways to see every column.

Keep outcome, control and operating measures separate.
DimensionEvidence to checkCommon misleading shortcut
Completed workRequired final state and quality criteria met for each eligible task.The model says the work is complete.
Authority and effectsAllowed reads and writes; no cross-scope access or prohibited effects.A useful result compensates for an unauthorized step.
Evidence and uncertaintyClaims supported; conflicts disclosed; missing evidence triggers the specified handover.A confident explanation is treated as verification.
RecoveryTimeout, restart and retry cases reconcile correctly without duplicate effects.Only the uninterrupted happy path is tested.
Human effortTime to inspect, correct, approve and resolve escalations.Generation time is presented as total time saved.
Operating performanceEnd-to-end latency and cost, including unsuccessful attempts and retries.Average model-call time and tokens per successful response.

A weighted dashboard can help prioritize improvements, but a high quality score should not erase a forbidden action. Treat defined prohibited effects as independent release blockers. OWASP’s excessive-agency analysis explains why permissions, available functions and execution freedom matter. An output-only grader cannot demonstrate that those controls held.

For useful-completion rate, divide completed eligible tasks by all eligible tasks presented, including those that timed out or were unnecessarily refused. Report correctly denied out-of-scope requests separately. For cost per useful completion, divide the included operating costs of the whole evaluated workload by useful completions. State whether human review and infrastructure are included; otherwise two cost figures may describe different economics.

Construct cases around decisions and failure boundaries#

Start with supported work drawn from domain experts, existing manual checks and known incidents. Keep a representative set for estimating everyday usefulness and a targeted challenge set for exposing failures. Do not blend them into a single ‘production success rate’: deliberately oversampling attacks and outages changes the question the score answers.

  • Normal: the correct record and required evidence are available, and the task should complete.
  • Ambiguous: two plausible records exist; the agent should clarify rather than guess.
  • Denied: the user lacks permission, including a case where an earlier permission has been revoked.
  • Adversarial: retrieved text attempts to redirect the task or request a prohibited action.
  • Uncertain effect: a write succeeds in the downstream system but its response is lost.
  • Interrupted: a run resumes after restart or receives approval after relevant inputs have changed.

Restore each trial’s environment: records, permissions, clock assumptions and any persistent memory. Record which dependencies are simulated and which use a test integration. A perfect mock cannot validate a real API’s timeout semantics. Test the relevant integration contract separately, without sending actual customer messages or modifying live customer records.

Keep the difficult cases intelligible. For every denial test, include a nearby permitted case so that ‘refuse everything’ cannot look successful. Retain a held-out set the team does not tune against; move newly discovered failures into regression coverage without pretending they remain unseen evidence.

Use code for facts, rubrics for judgement and people for calibration#

Use deterministic checks for record identity, amounts, recipients, policy results, duplicate effects and required receipts. Use a short rubric for qualities such as whether an explanation distinguishes confirmed evidence from inference. A model grader may help apply that rubric, but it cannot verify a database effect from a fluent completion message. Give it the relevant evidence and allow ‘insufficient evidence’ as a result.

Have domain reviewers independently assess a sample of passes, failures and uncertain cases. Compare their verdicts with the grader, inspect disagreements and version rubric changes. Hide candidate labels when comparing explanations where practical. Prevent the agent from editing the grading rules or expected outcomes; treat the trace supplied to a model grader as untrusted task data.

Avoid requiring one exact tool sequence unless that sequence encodes a real control, such as authorization before a protected read. There may be several valid routes to the same result. Conversely, checking only the final state can miss an unauthorized read. Outcome checks and boundary checks complement one another. This is the operational evaluation work developed in the Agentic AI Architecture & Security course, rather than a model leaderboard exercise.

Worked release decision: the higher score can still lose#

On a narrow screen, scroll the table sideways to see every column.

Hypothetical comparison on the same 100 eligible tasks; challenge cases evaluated separately.
MeasureWorkflow baselineAgent candidate
Useful completions82 of 10090 of 100
Total included operating costCHF 12CHF 18
Cost per useful completionAbout CHF 0.15CHF 0.20
Duplicate writes in challenge cases0 observed1 observed
Median human review time4 minutes3 minutes

The candidate improves observed completion and median review time in this example. It also fails the agreed no-duplicate-write gate. That is not a trade to hide inside a combined score. Investigate the retry path and rerun the affected and regression cases before granting write authority. A separately evaluated read-only draft mode may still be useful, but it is a different release scope with its own evidence.

Even after that defect is fixed, these totals alone do not establish a repeatable improvement. Inspect paired wins and losses, case mix, run-to-run variation and difficult subgroups. Report slow and expensive tails, not just medians. The eight extra completions may matter greatly, or they may be low-value cases whose review cost eliminates the benefit. Product owners must make that judgement against actual work.

Zero observed failures is not a zero failure rate#

Report how many distinct cases were tested and how many trials each received. Repeating one easy case one hundred times does not provide coverage of one hundred different situations. Likewise, allowing ten attempts and reporting whether any succeeded describes a different product experience from succeeding on the first attempt. Preserve both the attempt budget and the selection rule in the result.

For a simple numerical illustration, suppose 100 independent trials sampled from the same target distribution produce no failures. A one-sided 95% binomial upper confidence bound is 1 − 0.05^(1/100), approximately 2.95%. The calculation illustrates how much uncertainty remains; it does not predict a production failure rate. The assumptions often fail for correlated agent runs or a hand-picked challenge set. NIST’s confidence-interval reference explains the underlying binomial approach and why small failure counts need care.

Do not respond by running thousands of duplicate tests. Improve coverage of the actual decisions, permissions and dependency failures, then choose a sampling plan proportionate to the consequence of error. Statistical evidence cannot replace an enforceable permission boundary, and an adversarial test suite cannot certify that no untested attack exists.

A release gate your team can actually sign off#

A release gate is a recorded decision about a specific system and scope. Pin the application, model identifier, prompt, tool contracts, policy, relevant data/index versions and grading rules. Note any provider-managed component you cannot pin. The NIST AI RMF MEASURE function calls for documented evaluation under deployment-relevant conditions and explicit limitations; it does not supply a universal acceptable score.

Start the approved release within the scope you actually evaluated. Observe completion, interventions, unknown outcomes and effects after deployment, with sensitive content minimized. Define who can disable a capability and how in-flight work is reconciled. Re-evaluate when a model, policy, tool, source corpus or supported use case changes in a way that can alter behaviour. Monitoring is not a substitute for the initial gate; the initial gate is not a substitute for observing real use.

Sources and further reading

Primary sources checked on . Worked examples and worksheets are Ardevant teaching material; sources do not imply endorsement.

  1. Anthropic: Demystifying evals for AI agents

    Distinguishes task, trial, transcript and environmental outcome; discusses code, model and human graders and repeated-trial evaluation.

  2. NIST AI RMF: MEASURE function

    Supports documented test conditions, deployment-relevant evaluation, limitations and ongoing monitoring. It does not prescribe the illustrative release thresholds in this article.

  3. OWASP: LLM06:2025 Excessive Agency

    Supports evaluating the permissions, functionality and autonomy behind harmful effects, rather than relying on the wording of the final response.

  4. NIST/SEMATECH: Confidence intervals for proportions

    Explains confidence intervals and exact binomial methods. The zero-failure example here is an illustrative one-sided calculation under stated sampling assumptions.

Continue / Apply the method

Turn an agent demonstration into a release decision.

Work with Anton to define task outcomes, challenge cases and proportionate release gates for your architecture. Bring the current design and the failures you most need to prevent.

Author

Anton Selin

Software architect and your consulting partner.

Anton works on enterprise AI, cloud strategy, and agentic systems. Through Ardevant, he offers architecture consulting and practical training focused on the decisions your team needs to make.

About Anton and Ardevant