Back to Blog

What to Sample in an ISO 42001 Audit: An Evidence Field Guide

ISO 42001 requires "documented information" but doesn't specify what that looks like for AI. A field guide for what to sample in 2027-era audits.

Quick Answer

ISO 42001 requires "documented information" but doesn't specify what that looks like for AI. A field guide for what to sample in 2027-era audits.

ISO/IEC 42001:2023 requires "documented information" to evidence the operation of an AI management system. The standard does not specify what that documented information must look like for AI systems specifically. The first wave of certifications, starting in 2027, will establish the de facto evidence norms — and the firms doing those audits will be the ones setting them.

This is a field guide for what to sample in an ISO/IEC 42001 audit. It is organized by management system clause and by Annex A control category. It addresses the specific challenges that agentic AI introduces. And it includes a section on writing up findings for evidence quality issues — the category most likely to dominate the first audit cycles.

This guide assumes the engagement has been scoped per Part 2. It assumes the auditor is conducting a Stage 2 certification audit or an equivalent depth-of-procedure exercise. Readiness assessments use the same sampling categories but at lower coverage thresholds.

25–50
decisions per use case sampled for an agentic AI deployment over the audit period
4
evidence quality dimensions every sample should be evaluated against — integrity, completeness, linkage, retention
6
finding categories that will likely dominate the first audit cycle in 2027 — concentrated around impact assessment, foundation-model governance, lifecycle change records, affected-party infrastructure, agentic trace data, and AIMS-objective metrics

The sampling problem for AI audits

The unit of audit for ISO 42001 is the AI use case. The unit of evidence is more variable. For some clauses and controls, evidence is documentary — policies, procedures, records of management review. For others, evidence is operational — records of how the AI system actually behaved on specific inputs.

A defensible sampling plan addresses both dimensions:

  • Document sampling for the management system clauses (4-10): policies, scope statements, risk assessment outputs, management review minutes.
  • Operational sampling for the AI system itself: records of decisions, change events, evaluations, incidents.
What's Structurally Different About AI Sampling

Operational sampling for AI is structurally different from operational sampling for traditional IT systems. Where traditional IT sampling typically targets controls (e.g., "show me evidence of access reviews"), AI sampling often must target decisions (e.g., "show me what the system did for these specific inputs"). A control may operate correctly while individual decisions still produce harm. Sampling at the decision level is necessary for substantive findings on impact and affected-party requirements.

Sampling by management system clause

Clause 4 — Context of the organization

Document sampling: the organization's context statement, interested-party register, AI scope statement.

The auditor should verify that the interested-party register identifies affected individuals — not only commercial counterparties. Many context statements list customers, regulators, employees, and partners, but omit the end users of the AI itself. If the in-scope AI affects individuals (a loan applicant, a job candidate, a content moderator's subject), those individuals are interested parties under the standard.

The auditor should also verify that the AI scope statement enumerates use cases at a granularity that supports the audit. "We use AI" is not a scope statement; "we operate AI use cases UC-CS-01, UC-HR-04, and UC-RISK-12 with the boundaries described in [document]" is.

Clause 5 — Leadership

Document sampling: AI policy, AI roles and responsibilities matrix, evidence of leadership commitment (management review records).

The auditor should verify that AI roles include specific accountability for impact assessment and affected-party communications, not only technical AI ownership. Many organizations have well-defined technical ML ownership but no defined owner for the affected-party requirement; this is a finding pattern that will likely repeat in 2027 audits.

Clause 6.1.2 — AI risk assessment

This clause produces some of the most substantive audit work. The auditor should sample:

  • The risk assessment methodology document
  • The risk register, with attention to whether identified risks include both technical risks (model drift, prompt injection, training-data leakage) and societal or impact risks (bias, harm to affected parties, decision opacity)
  • The risk treatment plan and evidence of treatment implementation
  • Periodic reassessment records

Where the client claims to have integrated frameworks (e.g., DASF, OWASP LLM Top 10, MITRE ATLAS) into the risk identification process, the auditor should sample evidence of that integration — not just policy text claiming integration. A risk register that omits the OWASP LLM Top 10 categories the client claims to have considered is a finding.

Clause 6.1.4 and Annex A.8 — Impact assessment

The impact assessment is a separate requirement from risk assessment and is one of the most commonly under-evidenced controls.

Sampling should target:

  • The impact assessment methodology
  • Completed assessments for in-scope use cases, with attention to coverage of:
    • Affected parties (specifically named, not generic)
    • Types of harm (categorized, not just enumerated)
    • Severity gradations
    • Remediation plans where harm potential is identified
  • Periodic re-assessment evidence
  • Evidence that impact assessments inform deployment decisions, not just paper compliance

Where impact assessments are missing or thin, the audit finding should reference the specific affected-party population and the specific harm category not assessed.

Clause 7.5 — Documented information

The documented-information requirement underpins every other clause. The auditor should sample:

  • Document control records — who can edit what, when changes were made, version history
  • Retention policy and evidence of retention practice
  • Access controls on AI-related records

For AI systems specifically, the auditor should ask the client to demonstrate retrieval of a specific decision record from a specific date. Many production AI systems do not retain per-decision evidence at the granularity ISO 42001 requires; surfacing this gap during sampling produces a finding the client can act on.

Clause 8 — Operation

This clause produces sampling against the AI system lifecycle. The auditor should sample:

  • Records of AI system design decisions
  • Pre-deployment evaluations
  • Deployment approval records
  • Operational monitoring outputs
  • Change records (model version changes, prompt changes, configuration changes, training-data updates, retraining events)
  • Decommissioning records for any retired AI systems

The auditor should pay particular attention to evidence that changes are tracked at the right granularity. A change record that says "deployed v2 of the recommendation model" is less useful than a change record that says "deployed v2 of the recommendation model, with the evaluation summary at link X and the safety review at link Y." The former is a label; the latter is evidence.

Clause 9 — Performance evaluation

Sampling targets ongoing measurement of AI system performance against the management system's objectives.

The auditor should sample:

  • Performance metrics defined in the AIMS objectives
  • Records of metric collection
  • Trend analysis or threshold-based alerting evidence
  • Internal audit reports
  • Management review records

For metrics specifically, the auditor should verify that the metrics measured map to the objectives stated. Many clients measure operational metrics (latency, throughput, token usage) without measuring the metrics that bear on AIMS objectives (decision accuracy on protected populations, drift indicators, evaluation pass rates).

Clause 10 — Improvement

The improvement clause is often under-evidenced.

The auditor should sample:

  • Non-conformity records (from internal audits, customer complaints, incidents)
  • Corrective action records, with verification that the corrective action addressed the root cause
  • Continual improvement projects

Where the client has had AI-related incidents — a customer complaint about an AI-driven decision, a regulatory inquiry, a publicized model failure — the auditor should sample the incident record and the corrective action that followed.

Sampling by Annex A control category

The Annex A controls of ISO 42001 are grouped into nine categories. Sampling each category requires different evidence types.

A.5 / A.6 — Policies and organization

Document sampling: the AI policy itself, the organizational structure for AI governance, the responsibilities matrix. The auditor should verify that policies are signed, dated, and version-controlled. Unsigned or undated AI policies are surprisingly common.

A.7 — Resources

Document sampling: budget records, evidence of competency development (training, certifications), evidence that the organization has allocated appropriate resources to the AI management system. Where the AIMS budget is materially under-resourced relative to the AI footprint, this can support a finding under clause 7.1.

A.8 — Impact assessment

See clause 6.1.4 above. This Annex A control overlaps with the management system requirement; sampling can satisfy both.

A.9 — AI system lifecycle

See clause 8 above. Lifecycle controls are heavily evidence-dependent and benefit from per-use-case sampling.

A.10 — Third-party governance

Document sampling: vendor evaluation records, contracts (with attention to AI-specific clauses), ongoing monitoring outputs, evidence that vendor changes are tracked.

The Foundation-Model Evidence Gap

The auditor should pay particular attention to evidence of foundation-model provider evaluation — the most commonly under-evidenced sub-area of third-party governance. A client that can produce OpenAI's SOC 2 but not their own evaluation of OpenAI is not meeting the requirement. The standard requires evidence of the operator's evaluation, not the vendor's certification.

A.11 / A.12 / A.13 — Use of AI, third-party AI, information for interested parties

For B2C use cases or B2B use cases with affected-party requirements, the auditor should sample:

  • Affected-party communication artifacts (notifications, explanations, disclosures)
  • Records of affected-party requests and responses (e.g., GDPR Article 22 requests, equivalent rights under other regulations, internal disputes)
  • Per-decision explanation records, where the standard requires explanation

This sub-area is the most likely to produce findings related to evidence quality, because per-decision explanation records are infrastructure-intensive to produce.

Sampling agentic AI specifically

Compound or agentic AI use cases require sampling beyond the patterns above. The unit of evidence for these systems is the decision graph — the trace of which models, tools, and retrievers participated in producing a specific output, what inputs each received, what each produced, and how the components composed.

Sampling for agentic systems should:

  • Pull a sample of decisions — typically 25–50 for a use case covering the audit period
  • For each sampled decision, retrieve the full trace of contributing components (models, retrievers, tool calls, decision points, safety filters)
  • Verify that the trace is signed, time-stamped, and retained at audit-relevant time scales
  • Verify that the trace links to the controls catalog — the auditor should be able to navigate from a decision to the controls it operates under, and from a control to all decisions it claims to govern
  • Examine the trace for completeness: are all components instrumented, or are some operating opaquely?

Many production agentic systems in 2026 do not produce traces that satisfy these criteria. Findings related to agentic AI evidence sufficiency should be expected to be among the most common in 2027-era audits.

Evidence quality dimensions

Beyond sampling for specific clauses and controls, the auditor should evaluate four dimensions of evidence quality across the engagement.

Dimension
What to Verify
Weak Evidence Pattern
Integrity
Records are signed, hashed, or otherwise tamper-evident; protected against silent alteration. The SLSA / Sigstore families are useful reference points for what "tamper-evident" means operationally.
Records the client can edit silently from an admin console
Completeness
Evidence covers the population claimed; sampling-based completeness checks (drawing identifiers from independent sources) verify presence in the evidence store
Decision retention claimed for all customers but only a subset of customers have records
Linkage
Bidirectional linkage between controls and supporting evidence — auditor can navigate from a control to its evidence and from evidence back to the control
Evidence exists but cannot be linked from any control
Retention
Evidence covers the time period required by the management system's retention policy and applicable regulations
Observability-platform-derived evidence rolled off before the audit window opens

Common findings in 2027-era audits

Based on practice patterns visible in 2026 and the structural gaps in current GRC and AI infrastructure tooling, several finding categories are likely to dominate the first audit cycles:

  • Impact assessment thin or absent. Most clients have a risk assessment process; few have a separate impact assessment process. Expect this finding frequently in Stage 1 / readiness work.
  • Foundation-model governance under-evidenced. Clients can produce the foundation-model provider's SOC 2 but not their own evaluation of the provider.
  • Lifecycle evidence missing for change and decommissioning stages. Pre-deployment evidence is usually present; operational change records and decommissioning records are usually not.
  • Affected-party explanation infrastructure missing. Where the use case requires per-decision explanation, the infrastructure to retrieve specific decision records often does not exist at the required retention or granularity.
  • Per-decision evidence for agentic systems unavailable. The trace data exists in observability tools but is not retained at audit-relevant time scales and not linked to the controls catalogue.
  • Metrics measured but unmapped to AIMS objectives. Operational metrics tracked; AIMS-objective metrics not tracked or not visible to leadership.

Findings in these categories should be written to reference the specific clause or control and the specific evidence gap, not generic "AI governance is immature" language.

Writing findings on evidence quality

Findings on evidence quality (integrity, completeness, linkage, retention) require careful drafting. A finding that effectively says "the client's evidence is not good enough" without specificity will not survive client response.

The finding should:

Reference the specific clause or Annex A control

Name the exact requirement whose evidence is insufficient. "Annex A.10 third-party governance" — not "AI governance more broadly."

Reference the specific evidence quality dimension

Integrity, completeness, linkage, or retention. "The retention period of 30 days is insufficient against the lifetime-of-system requirement" is actionable. "Evidence is weak" is not.

Describe the specific evidence the auditor inspected

Document what was provided and what was missing. The finding should be reconstructable from the engagement file by a reviewer who wasn't on the engagement.

Identify what additional evidence would satisfy the requirement

The remediation path should be specific enough that the client can act on it without further auditor guidance.

Cite the standard's text or related guidance

Findings supported by specific clause text or recognized implementation guidance survive client challenge better than findings supported only by auditor judgment.

Well-drafted findings in this category produce useful Stage 2 remediation programs and protect the audit opinion from challenge. Generic findings produce client disputes, weaken the opinion, and damage the engagement relationship.

Sampling worksheet

For each in-scope use case, the engagement file should include a sampling worksheet covering:

  • Use case identifier and description
  • Deployment pattern (traditional ML / single foundation model / compound)
  • Sampled clauses and controls
  • Number of operational samples drawn
  • Evidence types collected per sample
  • Evidence quality assessment (per dimension above)
  • Findings identified

A standard worksheet template, used consistently across engagements, produces methodology defensibility and is the foundation for the kind of practice maturity that wins repeat audit business.

Conclusion

ISO 42001 sampling differs from ISO 27001 sampling in three structural ways: the unit of audit is the use case rather than the control, the operational evidence often must be sampled at the decision level rather than the control level, and the evidence quality dimensions — especially for agentic AI — are still being established. Firms that develop disciplined sampling plans now will set the norms the rest of the market follows.

In Part 4 of this series, we'll look at how the 2027-2030 AI audit market will differentiate, what positioning each kind of firm should consider, and what the talent and pricing trajectories look like.

Bring the evidence layer your audits assume exists

Most of the findings in this guide trace back to one root cause: AI infrastructure that wasn't built to produce audit-grade evidence. vCISO Lite is the integrated platform that closes that gap — signed records, controls linkage, retention discipline, and the bidirectional traceability that turns "we can't sample that" into a clean engagement file.

If you're building sampling discipline for the 2027 audit wave, or running readiness engagements that need to surface gaps before Stage 2, visit vcisolite.com to learn more and get started.

Where this matters next

Platform: APRI (AI-Powered Risk Intelligence)the verify-* MCP family — provenance + freshness + unverifiable_reasons envelopes that give ISO 42001 samplers something concrete to sample.

Platform: Compliancethe compliance surface that would house an ISO 42001 framework mapping if you're piloting one.

Where this matters next

The First AI Audits Hit in 2027. Most Mid-Market Companies Will Fail Them. — In 18 months, a new generation of audit opinions will start landing in mid-market boardrooms

EU AI Act Article 12: What AI Logging Requirements Mean for Audit Firms and Their Clients — EU AI Act Article 12 is already applicable to high-risk AI systems newly placed on the EU market

The 2027-2030 AI Audit Market: How Assurance Firms Will Differentiate — By 2030, the firms that started early will have shaped how AI assurance is delivered

How often is your compliance AI actually right? — Every vendor pitching an 'AI compliance agent' makes the same promise — set it loose on your controls and it will attest while…

Share this article:

Ready to build your security program?

See how easy it can be.