Back to Blog

How often is your compliance AI actually right?

Every vendor pitching an 'AI compliance agent' makes the same promise — set it loose on your controls and it will attest while you sleep. The honest answer to 'how often is it right' is that nobody knows, and the market has organized itself around not having to say so. Here is what a credible answer looks like, and what it would take to publish one.

Quick Answer

Every vendor pitching an 'AI compliance agent' makes the same promise — set it loose on your controls and it will attest while you sleep. The honest answer to 'how often is it right' is that nobody knows, and the market has organized itself around not having to say so. Here is what a credible answer looks like, and what it would take to publish one.

Every vendor pitching an "AI compliance agent" right now is making one version of the same promise. Set it loose on your controls, and it will attest, triage, and remediate while you sleep.

Here is the question the pitches do not answer. How often is the agent right? Not on a demo. Not on a benchmark you have never seen. On real controls, when nobody is watching.

The honest answer, today, is nobody knows— and the market has organized itself around not having to say so.

What the market actually delivers

Independent reporting on production AI-SOC and AI-compliance deployments in 2026 finds a persistent gap between the marketing and the running software. The agents are doing enrichment, summarization, drafting, and non-judgment workflow. Humans are still making the calls that carry liability.

0
Production deployments observed taking consequential action without a human approval gate
Help Net Security, March 2026
HITL
Posture Vanta states for its new compliance agents — humans-in-the-loop retained for final decisions
Vanta product announcement, March 2026
Bounded
Autonomy framing for CrowdStrike's Charlotte Agentic SOAR — runs under analyst-defined guardrails
CrowdStrike Charlotte AI, November 2025

Read the actual product copy from the leading vendors and the language is consistent. Suggest-only. Humans retain final decision authority. Bounded by analyst command. These are not admissions of weakness. They are the floor of what is responsible to ship when no one can prove the agent is correct enough to act on its own.

The market position, plainly

Every credible vendor in compliance AI has converged on the same answer: the agent helps a human decide; the human carries the call. The agents that claim more either gate "more" on different language ("approve once, runs forever") or are not in production at scale.

Why the gap exists — it isn't capability

Connecting a capable model to a tool interface is routine engineering in 2026. Anyone with two engineers and three months can ship an agent that can read a configuration, judge it against a control, and write the action.

The unsolved problem is the one nobody wants to put a number on. How do you provethe agent is trustworthy enough to be permitted to act — to an auditor, to a board, to the regulator that will eventually ask?

That is a measurement problem, and the measurement does not exist yet. Every vendor claim about agent accuracy I have seen in this category falls into one of three buckets: a demo, a self-reported score on a public benchmark, or no number at all. The first two are unsafe for the decision; the third is at least honest.

What "accuracy" should mean for a compliance agent

In compliance the cost of the two ways an AI can be wrong is wildly asymmetric. A blended "accuracy" number averages them and hides the one that matters.

False positive
False negative
What happened
Agent flagged a passing control as FAIL
Agent attested a failing control as PASS
Who pays
An analyst, in time and trust
The company, in liability
When it surfaces
Immediately — the alarm goes off
Later — when the audit, breach, or incident exposes it
How to recover
Tune the agent and move on
There is no recovery; the attestation is in the record

The number that should determine whether a compliance agent is allowed to act on its own is the wrong-attestation rate— the rate at which the agent says PASS when the truth is FAIL. In the statistical literature this is the false-negative rate. In plain English: how often does the AI quietly approve something that is actually broken?

That number, on its own, decides whether the agent earns autonomy. A compliance agent with 98% blended accuracy and a 12% wrong-attestation rate is not a good agent. It is a fast way to attest broken controls.

Why existing AI benchmarks will not save us

The temptation, of course, is to publish a benchmark and let scores speak for themselves. The AI industry has tried this. It has not gone well.

The cautionary tale: SWE-bench

SWE-bench was the dominant agent-evaluation benchmark for code-fixing agents in 2024 and 2025. Then, in early 2026, an adversarial-strength re-evaluation found the top system's success rate fell from 78.8% to 62.2% under a stricter rubric — and that roughly one in five "solved" instances were semantically incorrect (Yu et al., arXiv 2603.00520, 2026). A month later, OpenAI publicly deprecated SWE-bench Verified after finding that training data contamination was inflating scores (OpenAI, February 2026). Scale AI's response was to build a successor on private, legally-inaccessible codebases — which reported materially lower, more honest numbers (Scale AI SWE-Bench Pro, 2025).

Three failure modes turned the industry's most-used coding-agent benchmark into a stat nobody trusts:

  1. Under-discrimination. The rubric was too loose; the score did not separate good agents from less-good ones.
  2. Training contamination. The test data leaked into the models being tested.
  3. Vendor self-report. The scores were reported by the people graded by them.

Any compliance-AI benchmark that does not engineer against all three from day one will repeat the same arc — and the stakes are materially higher. A wrong SWE-bench score embarrasses a lab. A wrong compliance-attestation score is a regulatory exposure.

What a credible compliance-AI benchmark looks like

A benchmark engineered against the three failure modes above looks different from anything currently in the market. Five non-negotiables:

Two-tier corpus — public for calibration, private for scoring

The public tier — adjudicated cases sourced from regulator orders, consent decrees, audit opinions, breach post-mortems — exists so anyone can see how the test is constructed and how a particular verdict was scored. Calibration and credibility only. The private tier — auditor-graded cases, never published — is the only set that produces a reported score. Public sets are assumed contaminable. Private sets are the only ones that can certify.

Contamination resistance, certified

Every private case is constructed after the evaluated model's training cutoff, with chain-of-custody to prove it, and never enters any trainable store. A standard membership-inference probe confirms the private set looks like random data to the model — AUC near 0.5. This is the lesson from SWE-Bench Pro: if it is on the open internet, assume it is in training.

Stratified reporting, never a blended number

Wrong-attestation rates are reported per framework, per control family, and split between cases with explicit adjudication versus cases reconstructed from public records. A blended number hides the failure modes that matter — a single weak control family is enough to disqualify autonomy on that family even when the average looks fine.

Certify on the upper bound, not the point estimate

A point-estimate accuracy of 92% on 600 cases is not, statistically, the same as a guarantee that the true wrong-attestation rate is at most 8%. The honest claim is a confidence interval. The certifying number should be the upper bound of the 95% confidence interval — conservative by construction. Finite-sample uncertainty withholds autonomy rather than granting it optimistically.

Validate the benchmark, not just the agent

Run the benchmark against itself. Re-score under a strengthened rubric and report the discrimination gap. Check inter-rater reliability of the gold set — if the credentialed graders cannot agree, the gold is not gold. Run the contamination probe. Ablate the runtime monitor and report the uplift it actually buys. A certification number that comes out of a benchmark that has not been validated is itself untrustworthy.

The Trustworthy Autonomy approach

This is the methodology we are publishing under the brand Trustworthy Autonomy. The full specification covers five measured trust components — decision accuracy, reproducibility, runtime reliability, provenance completeness, and adversarial robustness — combined through a per-action-category decision rule that grants autonomy only when every component clears a pre-registered threshold with stated confidence.

The four validation tests above run on the benchmark itself, every version, and the results are published alongside the certification. The accuracy claim is conservative on purpose. The thresholds are a risk-appetite policy, not a vendor-friendly constant, and they are public.

The runtime monitor — the second of two gates, the one that watches a run as it executes rather than a build before it ships — is treated as a distinct measurable property, not as a decorative claim. We measure its detection recall against injected failures and report the strategic-quality uplift it buys via ablation. A monitor that does not move the number does not count.

The right to act autonomously is not a marketing claim. It is a number that has to be earned, measured, and proven — and the test that produced the number has to be one a third party can pick apart.

Where this stands today, honestly

The Trustworthy Autonomyframework specification and the benchmark methodology are public. The worked example in the methodology spec uses illustrative figures and deliberately ends in a denial — the conservative bound on the wrong-attestation rate for SOC 2 CC6 sits above the threshold — precisely to show how the test behaves when the evidence does not yet support autonomy.

What we are doing now is building the floor a real certification needs. The public adjudicated-case corpus — sourced from regulator orders, audit opinions, and breach post-mortems — is being seeded. The private auditor-graded scoring set is being sized against the statistical bound that determines whether any verdict on it is meaningful. Independent administration of the scoring set is the next milestone after that.

When the first real certification lands, it will be of a specific candidate build, on a specific action category, on a held-out set that survived the contamination probe. It will publish every component measurement, the binding gates, the verdict, and the remediation path if denied. Until then the framework is the answer to the buyer's question — not a score on a single agent.

If you are buying compliance AI right now and a vendor cannot answer the question at the top of this article with a number, an interval, and a methodology a third party could attack, the responsible move is to keep the human in the loop. The rest of this series unpacks what "in the loop" should look like, where it should give way, and what the substrate underneath all of it has to be to make any of it provable.

Where this matters next

The trust ladder: from approve-everything to goal-orientedwhat graduated autonomy looks like in practice, with the same public S3 bucket triggering a SOC 2 drift event handled three different ways across the tiers, and why human-in-the-loop is an authority control and not a substitute for the runtime monitor.

APRI: AI-Powered Risk Intelligencethe compliance-AI surface built to the standards this article calls for — measured, cited on every claim, audited on an independently verifiable chain.

Trustworthy Autonomyour POV on autonomous compliance — the tiered trust ladder, the evaluation regime, and what evidence would need to land before any vendor could cross the bar.

Share this article:

Ready to build your security program?

See how easy it can be.