AI Audit Evidence Checklist for Internal Audit and Compliance Teams

Share Article

Table of Contents

Spending on AI governance platforms will reach $492 million in 2026 and pass $1 billion by 2030, according to a February 2026 Gartner press release, which also found that organizations running dedicated governance tooling are 3.4 times more likely to govern AI effectively than those working from spreadsheets. Behind that figure sits a quieter problem. When an internal auditor or a certification body asks an organization to prove its AI is governed, most teams find their evidence scattered across drives, undated and impossible to trace back to a control. This checklist closes that gap. It lays out the 14 evidence artifacts an AI audit expects, maps each one to its source clause in ISO/IEC 42001, the EU AI Act and the NIST AI Risk Management Framework and shows what internal auditors actually test once your records are on the table.

What counts as audit evidence in an AI management system

Audit evidence is anything that lets a third party confirm a control operated as intended, without taking your word for it. Inside an AI management system, that evidence falls into four categories and auditors weigh them differently.

Documented information comes first: policies, procedures, the risk assessment methodology, and the AI objectives leadership signed off on. ISO/IEC 42001 Clause 7.5 governs this category and carries a demand most teams overlook. Documented information has to be controlled, meaning version-tracked, dated and tied to an owner. A policy with no revision history is weak evidence even when its content is correct.

Records are the second type, the outputs that prove a process ran: a completed risk assessment, a signed impact assessment, a management review minute, a corrective action log. Where documented information states what should happen, records prove that it did.

System-generated evidence is the third type, and increasingly the most scrutinized. Article 12 of the EU AI Act requires high-risk systems to keep automatic logs of their operation across their lifetime, and Article 19 makes those logs retainable and producible on request. For US teams the expectation is the same under NIST AI RMF’s MEASURE function: model performance metrics, drift monitoring, access logs and inference records.

Attestations round out the set: sign-offs, approvals, and competence records that connect a decision to a named, qualified person. An auditor reading an approval wants to know who gave it, when and whether they were authorized to.

The AI audit evidence checklist: 14 artifacts auditors expect

Start with this list. Each item is an artifact an auditor can request by name, grouped by the governance function it belongs to. The framework references in brackets tell you which clause or article the artifact satisfies, so you can defend why it exists.

Governance foundation

  1. AI policy and objectives. The board-approved AI policy plus measurable governance objectives. Proves leadership commitment. [ISO 42001 Clause 5.2, 6.2]
  2. Roles and responsibilities matrix. Who owns each AI control, including the accountable executive. [Clause 5.3; NIST GOVERN 2]
  3. AI system inventory. A complete register of AI systems in use, each with a risk classification, owner, and lifecycle status. This is the single most requested artifact in an AI audit, and the one most often missing or incomplete. [Clause 8.1; EU AI Act Art. 49; NIST MAP 1]

Risk and impact

  1. Risk assessment records. The methodology plus a completed assessment for each system, scored and dated. [Clause 6.1.2, 8.2; EU AI Act Art. 9]
  2. Impact assessments. Assessments of effects on individuals and groups, including a fundamental rights impact assessment where EU AI Act Article 27 applies to your deployment. [Clause 6.1.4; Annex A.5]
  3. Risk treatment plan and Statement of Applicability. Which Annex A controls apply, which are excluded, and the justification for each exclusion. [Clause 6.1.3]

Data and technical

  1. Data governance records. Provenance, quality checks, bias testing, and the lawful basis for training and input data. [Annex A.7; EU AI Act Art. 10]
  2. Technical documentation. Model design, intended purpose, performance characteristics, and known limitations for each system. [Clause 7.5; EU AI Act Art. 11 and Annex IV]
  3. Testing and validation evidence. Accuracy, robustness, and security test results, captured before deployment and after material change. [EU AI Act Art. 15; NIST MEASURE 2]

Operation and oversight

  1. Event logs and audit trails. Automatically generated operational logs that reconstruct what a system did and when. [EU AI Act Art. 12 and 19]
  2. Human oversight records. Evidence that designated people can monitor, intervene in, and override the system, with logs showing they did when needed. [Annex A.9; EU AI Act Art. 14]
  3. Change and approval records. Who approved each model version, deployment, or material change, with dates and authority. [Clause 8.1; Annex A.6]

Monitoring and assurance

  1. Monitoring and incident records. Performance monitoring, drift detection, incident logs, and the corrective actions that followed. [Clause 10.2; EU AI Act Art. 72 and 73]
  2. Internal audit and management review outputs. The internal audit reports and management review minutes that close the governance loop. [Clause 9.2, 9.3; NIST GOVERN 4]

Your audit evidence starts with a complete AI system inventory that identifies every AI model, tool, agent, and vendor in use.

PRO TIP Two artifacts sit outside this numbered set but get requested constantly: third-party and supplier records for any procured AI [Annex A.10; EU AI Act Art. 25], and competence and training records that show the people running your AI are qualified [Clause 7.2; EU AI Act Art. 4]. Add both to your evidence pack before the auditor asks.

Cross-framework evidence mapping: ISO 42001, EU AI Act, NIST AI RMF

Most teams keep three separate compliance binders, one per framework, and duplicate effort across all of them. The reality is that a single well-built evidence artifact usually satisfies all three at once, because the frameworks converge on the same underlying controls. Build the evidence once, map it three ways. This table does the mapping.

Evidence artifactISO/IEC 42001:2023EU AI ActNIST AI RMF 1.0
AI policy & objectivesClause 5.2, 6.2; Annex A.2Art. 17 (QMS)GOVERN 1.1, 1.2
AI system inventoryClause 8.1; Annex A.4Art. 49, 71 (registration)MAP 1.1, 1.6
Risk assessment recordsClause 6.1.2, 8.2Art. 9 (risk management)MAP 5; MEASURE 2
Impact assessmentClause 6.1.4; Annex A.5Art. 27 (FRIA)MAP 1, 3
Data governance recordsAnnex A.7Art. 10 (data governance)MAP 2; MEASURE 2.11
Technical documentationClause 7.5; Annex A.6Art. 11 + Annex IVMAP 4
Testing & validationAnnex A.6Art. 15 (accuracy, robustness)MEASURE 2
Event logs / audit trailAnnex A.6Art. 12, 19 (logging)MEASURE 2; MANAGE 4.1
Human oversight recordsAnnex A.9Art. 14 (human oversight)GOVERN 2; MANAGE 2
Incident & corrective actionClause 10.2Art. 73 (serious incidents)MANAGE 4
Supplier / third-party recordsAnnex A.10Art. 25 (value chain)GOVERN 6; MAP 4.1
Internal audit & reviewClause 9.2, 9.3Art. 17 (QMS checks)GOVERN 4

Every evidence package should include a documented AI risk assessment to demonstrate how risks were identified and treated.

Note on EU AI Act timing. Under the Digital Omnibus provisional agreement reached on 7 May 2026, obligations for standalone high-risk systems in Annex III are deferred to 2 December 2027 and for product-embedded systems to 2 August 2028, while Article 50 transparency duties and enforcement powers remain live from 2 August 2026 (see the Council statement). These dates are provisional until the text is published in the Official Journal.

Evidence by AI lifecycle stage

A document list tells you what to collect. It does not tell you when each piece should have been created, and that is exactly what a sharp auditor probes. ISO/IEC 42001 Annex A.6 frames governance around the AI system life cycle for this reason: evidence is strongest when it shows a control operated at the stage where it mattered, not when it was assembled the week before the audit.

Lifecycle stageEvidence the stage should produce
Planning & designIntended purpose statement, initial risk classification, design-stage impact assessment
Data & developmentData provenance and quality records, bias testing, training documentation
ValidationAccuracy, robustness, and security test results against defined acceptance criteria
DeploymentRelease approval, deployment configuration, human oversight design, go-live sign-off
Operation & monitoringOperational logs, performance and drift monitoring, incident and corrective action records
RetirementDecommission record, data disposal evidence, retention of documentation per policy

Mapping evidence to the lifecycle does two things. It exposes the stage where your records thin out, usually monitoring and retirement, and it lets you answer the auditor’s real question: did this control operate continuously, or only once?

What internal auditors actually test

Handing an auditor a folder of documents is not the same as passing an audit. Internal auditors apply the same evidence-gathering techniques to AI that they apply everywhere else, drawn from established internal audit practice. Knowing the four techniques tells you how robust each artifact needs to be.

Inquiry. The auditor asks how a control works. Useful for context, but inquiry alone is the weakest form of evidence. Expect every verbal answer to be tested against a record.

Observation and walkthrough. The auditor watches the control operate, or traces one AI system end to end: from its entry in the inventory, through its risk assessment, into its monitoring logs. A walkthrough exposes whether your artifacts actually connect to each other.

Inspection of records. The auditor reads the evidence itself and checks dates, owners, approvals, and version history. This is where uncontrolled documents fail.

Reperformance. The auditor re-runs part of the control, for example recalculating a risk score or re-querying a log, to confirm it produces what you claimed. The IIA’s AI auditing guidance encourages this kind of independent testing for AI-specific risks such as drift and bias.

Sampling sits across all four. Auditors rarely review every system. They pick a sample, often weighted toward high-risk systems, and test it deeply. If your high-risk AI is well evidenced and your low-risk AI is not, you are still exposed, because the auditor chooses the sample, not you.

Evidence gaps that quietly trigger findings

Most audit findings are not caused by missing controls. They are caused by controls that worked but were poorly evidenced. These are the recurring gaps.

  • Spreadsheet inventories with no version history. The AI register exists, but no one can show what it looked like six months ago, so the auditor cannot confirm it was maintained.
  • Undated or unsigned approvals. A model went live with sign-off, but the approval has no date or no named approver, which makes it unverifiable.
  • Risk assessments that do not tie to controls. Risks are documented, but nothing links them to the Annex A controls or treatment decisions meant to address them.
  • Shadow AI with no logs. Systems adopted by teams outside the governance process leave no audit trail at all, and auditors increasingly ask how you find them.
  • Point-in-time evidence for period-of-time controls. A single screenshot proves a control existed once. Continuous monitoring evidence proves it operated throughout the audit period, which is what certification requires.

Every gap on this list shares a root cause: evidence created as a one-off rather than generated as a byproduct of the control running. Close the gap by making evidence automatic, not manual.

How long to keep AI audit evidence and who owns it

Two questions decide whether your evidence survives contact with an audit: how long you keep it, and who is accountable for it.

Retention. The EU AI Act sets the longest clock. Article 18 requires providers of high-risk systems to keep technical documentation, quality management records, and related evidence for 10 years after the system is placed on the market or put into service. Article 19 requires automatically generated logs to be retained for an appropriate period of at least six months unless other law sets longer. ISO/IEC 42001 does not fix a number; it requires you to define retention in policy and then meet it, which means an undefined retention period is itself a nonconformity.

Ownership. Every artifact needs a named owner who is accountable for keeping it current, not just whoever happened to create it. The cleanest way to assign this is a responsibility matrix tied to the inventory: each system has an owner, and each evidence type has a custodian. When ownership is implicit, evidence ages out and no one notices until the auditor does.

From spreadsheet to system: managing evidence at audit scale

A spreadsheet handles your first AI audit. It struggles at the second and fails at the third, when the auditor asks for version history, continuous monitoring evidence, and proof that every system in the inventory maps to its controls. The volume is the problem: a mid-size enterprise governing a few dozen AI systems across three frameworks is tracking hundreds of evidence artifacts, each with its own owner, retention clock, and lifecycle stage.

This is the work an AI governance platform is built to absorb. Govern365.ai‘s AI model registry maintains a live inventory where each system is automatically mapped to its applicable ISO 42001 clauses, EU AI Act obligations, and NIST AI RMF functions, so the cross-framework mapping in this article is generated rather than maintained by hand. Its audit evidence management captures records, logs, and approvals as the controls run, with the version history and ownership that uncontrolled documents lack.

The payoff is not just a tidier audit. It is the difference between assembling evidence under deadline pressure and producing it on demand, which is the position every internal audit and compliance team wants to be in before the certification body arrives.

Frequently asked questions

What is AI audit evidence?

AI audit evidence is any record that lets an auditor confirm an AI control operated as intended, independent of your assurances. It spans four types: documented information such as policies, records such as completed risk assessments, system-generated logs, and attestations such as dated approvals. ISO/IEC 42001 Clause 7.5 requires this evidence to be controlled, meaning versioned, dated, and owned.

What evidence does an ISO 42001 audit require?

An ISO/IEC 42001 audit requires evidence that your AI management system was planned, operated, and improved as documented. Core artifacts include the AI policy, the system inventory, risk and impact assessments, the Statement of Applicability, operational logs, internal audit reports, and management review minutes. The auditor tests whether these connect and whether controls operated across the full audit period, not just once.

How is AI audit evidence different from SOC 2 or ISO 27001 evidence?

The evidence techniques are the same, but AI adds artifact types that security audits do not cover: model impact assessments, bias and data governance records, drift monitoring, and human oversight logs. AI evidence also has to address how a model behaves over time, so point-in-time proof is rarely enough. Continuous monitoring evidence matters far more than in a traditional controls audit.

How long do I need to keep AI audit evidence?

Under EU AI Act Article 18, providers of high-risk systems keep technical documentation and quality management records for 10 years after the system is placed on the market. Automatically generated logs are kept for at least six months under Article 19. ISO/IEC 42001 sets no fixed period but requires you to define retention in policy and meet it, so an undefined period is a finding in itself.

Who is responsible for collecting AI audit evidence?

Accountability sits with the AI governance or compliance function, but ownership of individual artifacts should be distributed. Each AI system needs a named owner, and each evidence type needs a custodian, usually mapped in a responsibility matrix tied to the inventory. Internal audit then tests the evidence independently rather than collecting it, to preserve objectivity.

Does the NIST AI RMF require audit evidence?

The NIST AI Risk Management Framework is voluntary and prescribes no certification audit, so it does not mandate evidence the way ISO does. In practice its MEASURE and MANAGE functions expect the same artifacts: documented risk analysis, performance measurement, monitoring, and incident response. US organizations increasingly evidence NIST alignment because state laws and customers reference the framework as a benchmark.

What is the most common reason organizations fail an AI audit?

The most common cause is not absent controls but poorly evidenced ones: undated approvals, spreadsheet inventories with no history, and risk assessments that do not link to controls. The second is point-in-time evidence where the auditor needs proof a control operated continuously. Both are fixed by generating evidence automatically as controls run, rather than assembling it before the audit.

Closing the evidence gap

An AI audit is won or lost on evidence, not intentions. The teams that pass are not the ones with the most controls; they are the ones whose controls leave a trail an auditor can follow, from the system inventory through risk, oversight, and monitoring, each artifact dated, owned, and tied to a framework clause. Start by building the 14-artifact checklist above and mapping each item to the clauses it satisfies, then check where your evidence is created as a one-off rather than generated automatically. Those one-off points are your audit risk.

When you are ready to make evidence produce itself, start your 14-day free trial of Govern365.ai, by the Global AI Certification Council, and see your AI systems mapped to ISO 42001, the EU AI Act, and NIST AI RMF from day one.

Stay ahead of the curve

Join 5,000+ industry leaders who receive our weekly briefing on AI governance and secure enterprise collaboration.

About the Author

Dr Faiz Rasool

Director at the Global AI Certification Council (GAICC) and PM Training School

Globally certified instructor in ISO/IEC, PMI®, TOGAF®, and Scrum.org disciplines with hands-on experience in ISO/IEC 42001 AI governance across the US, EU, and Asia-Pacific.

Summarize with AI

AI-Powered Data Governance Platform

Secure, Govern, and Collaborate on Sensitive Data—All Within Microsoft 365

Further Reading

Related Insights

eu-ai-act-us-companies-applicability-records-controls

EU AI Act for US Companies: Applicability, Records and Controls

Spending on AI governance platforms is projected to reach $492 million in 2026 and surpass

Read More →
ai-governance-roadmap-mid-market-risk-teams

US AI Governance Roadmap for Mid-Market Risk Teams

Forty-five state legislatures introduced more than 1,561 AI-related bills by March 2026 alone, according to

Read More →
us-state-ai-law-tracker-compliance-teams

US State AI Law Tracker: What Compliance Teams Must Know Now

State lawmakers introduced 1,561 AI-related bills across 45 states in the first quarter of 2026

Read More →

Summarize with AI

Transforming AI Risks into Strategic Assets.

Request a Personalized Demo

Our governance experts will walk you through the platform and help you map out your ISO 42001 or EU AI Act roadmap.