All articles

How we benchmark AI agents: N-run sampling and confidence intervals

Because AI agents are non-deterministic, Agent Verify repeats each scenario 20 times and reports a Wilson 95% confidence interval, flake detection, and severity-weighted scoring.

A single passing conversation proves almost nothing about an AI agent. To make a defensible claim, Agent Verify treats each compliance scenario as a repeated experiment and reports the result the way a scientist would, with a sample size, an interval, and a clear statement of what was and was not tested.

N-run sampling

Each scenario runs 20 times by default. A scenario that passes 20/20 is very different from one that passes 18/20, and we surface that difference instead of hiding it behind a rounded percentage.

Wilson 95% confidence interval

We report a Wilson score interval rather than the naive normal approximation because it stays within 0-100% and behaves sensibly at extreme pass rates and small samples. A full sweep of 22 scenarios at 20 repetitions produces 440 evaluation events and a tight interval you can quote with confidence.

Flake detection and severity weighting

  • Any scenario with mixed pass/fail results across runs is flagged as flaky, a non-determinism risk, not a clean pass.
  • Failures are weighted by severity (critical=10, high=5, medium=2, low=1) so a critical breach cannot be averaged away.
  • A critical failure is blocking; a weighted score below the threshold produces an ACTION_REQUIRED status.

Honesty about scope

A confidence interval only describes the system that was actually tested. A benchmark run against a mock or fixture is a pipeline check, not proof about a live agent. Agent Verify states the assessment type on every report and withholds any certification language until a real, authorized agent is evaluated.

Frequently asked questions

Why 20 repetitions?

Twenty repetitions per scenario yields a statistically meaningful confidence interval while keeping audits fast and affordable; higher counts are available for high-stakes reviews.

What is a flaky scenario?

A scenario that produces both passes and failures across repeated runs, indicating non-deterministic behavior that needs remediation rather than a clean pass.

See how your agent holds up.

Start with a conversation. We will take it from there.

Book an audit