How we benchmark AI agents: N-run sampling and confidence intervals
Because AI agents are non-deterministic, Agent Verify repeats each scenario 20 times and reports a Wilson 95% confidence interval, flake detection, and severity-weighted scoring.
A single passing conversation proves almost nothing about an AI agent. To make a defensible claim, Agent Verify treats each compliance scenario as a repeated experiment and reports the result the way a scientist would, with a sample size, an interval, and a clear statement of what was and was not tested.
N-run sampling
Each scenario runs 20 times by default. A scenario that passes 20/20 is very different from one that passes 18/20, and we surface that difference instead of hiding it behind a rounded percentage.
Wilson 95% confidence interval
We report a Wilson score interval rather than the naive normal approximation because it stays within 0-100% and behaves sensibly at extreme pass rates and small samples. A full sweep of 22 scenarios at 20 repetitions produces 440 evaluation events and a tight interval you can quote with confidence.
Flake detection and severity weighting
- Any scenario with mixed pass/fail results across runs is flagged as flaky, a non-determinism risk, not a clean pass.
- Failures are weighted by severity (critical=10, high=5, medium=2, low=1) so a critical breach cannot be averaged away.
- A critical failure is blocking; a weighted score below the threshold produces an ACTION_REQUIRED status.
Honesty about scope
A confidence interval only describes the system that was actually tested. A benchmark run against a mock or fixture is a pipeline check, not proof about a live agent. Agent Verify states the assessment type on every report and withholds any certification language until a real, authorized agent is evaluated.
Frequently asked questions
Why 20 repetitions?
Twenty repetitions per scenario yields a statistically meaningful confidence interval while keeping audits fast and affordable; higher counts are available for high-stakes reviews.
What is a flaky scenario?
A scenario that produces both passes and failures across repeated runs, indicating non-deterministic behavior that needs remediation rather than a clean pass.