AI compliance evidence is the set of dated, attributable records that show an AI obligation was met: the technical file, the event logs, the risk assessment, the approval, the incident report. Auditors and market surveillance authorities test records rather than intentions. Each obligation in the EU AI Act, ISO/IEC 42001 and the NIST AI Risk Management Framework has a specific artefact that answers it, and a governance programme is judged on whether it can produce that artefact on request.
Evidence gaps show up in the first audit, not before it. ISACA published its 2026 AI Pulse Poll on 5 May 2026, drawn from more than 3,400 digital trust professionals. In it, 39 percent said they did not know whether their organisation had a documented process for shutting down or overriding an AI system, and 56 percent did not know how long halting a system would take. Neither figure describes a policy problem. Both describe a records problem.
An auditor tests records; a policy only states intent
Governance decks tend to merge two things an auditor keeps apart. A policy, a standard or a framework sets out what an organisation intends to do. Audit evidence records what it did, when, and on whose authority.
ISACA’s poll puts numbers on the first half: 38 percent of organisations hold a formal AI policy, up from 28 percent a year earlier, while 25 percent hold none. Policy coverage is improving. Record coverage is a separate track, because a policy sitting on a shared drive still generates no approval, no test result and no log.
Practitioners phrase the test more bluntly. A version of this question circulates on r/grc: if a regulator asked you to prove your AI was safe yesterday, what would you hand over?
Three record families answer three different questions
Chapter III of Regulation (EU) 2024/1689 asks for three kinds of record rather than one document. Treating them as one file is the most common design error in an AI governance programme, because each answers a question the others cannot.
What you built sits in the technical documentation. Article 11 requires that file to be “drawn up before that system is placed on the market or put into service”, with its contents set out in Annex IV: intended purpose, architecture, training data characteristics, test results, known limitations.
What the system did sits in the logs. Article 12(1) requires high-risk systems to “technically allow for the automatic recording of events (logs) over the lifetime of the system”. Article 12(3) goes further for remote biometric identification and requires the log to capture “the identification of the natural persons involved in the verification of the results”. Field-level detail on that specification sits in what Article 12 logging has to capture.
What happened after launch sits in monitoring and incident records. Article 72(1) requires providers to “establish and document a post-market monitoring system”, and Article 72(3) states that “the post-market monitoring plan shall be part of the technical documentation referred to in Annex IV”, which pulls the monitoring design back into the file built before launch.
The obligation-to-evidence map pairs every duty with its record
Evidence attaches to duties one row at a time. Each row below names the obligation, the instrument and reference behind it, the artefact that satisfies it, the role that usually owns it, and the clock attached. Rows follow the order an audit tends to run: build first, operate second, oversee third.
| Obligation | Instrument and reference | Record that answers it | Typical owner | Retention |
| Risk management system | EU AI Act Article 9 | Risk assessment and AI risk register, dated and versioned | AI governance manager | Travels with the technical file |
| Data and data governance | EU AI Act Article 10(2) | Dataset documentation: design choices, provenance, bias examination and mitigation | Data owner | Travels with the technical file |
| Technical documentation | EU AI Act Article 11, Annex IV | The technical file, complete before launch | Product or engineering lead | 10 years (Article 18(1)(a)) |
| Record-keeping | EU AI Act Article 12 | Event logs, including reviewer identity for biometric systems under Article 12(3) | System owner | At least six months (Articles 19 and 26(6)) |
| Transparency to deployers | EU AI Act Article 13 | Instructions for use, issued and version-controlled | Product lead | Travels with the technical file |
| Human oversight | EU AI Act Article 14(4) | Oversight design note plus records of override and non-use decisions | Business process owner | Logs at least six months |
| Quality management system | EU AI Act Article 17(1) | Regulatory compliance strategy, design control and quality assurance procedures | Quality or compliance lead | 10 years (Article 18(1)(b)) |
| Conformity assessment | EU AI Act Article 43 | Assessment record and notified body decisions where applicable | Compliance lead | 10 years (Article 18(1)(c) and (d)) |
| Declaration of conformity and marking | EU AI Act Articles 47 and 48 | EU declaration of conformity, CE marking file | Compliance lead | 10 years (Article 18(1)(e)) |
| Registration | EU AI Act Article 49 | EU database entry made before market placement | Compliance lead | Kept current for the life of the system |
| Post-market monitoring | EU AI Act Article 72 | Monitoring plan inside Annex IV, plus the data collected against it | System owner | Plan follows the technical file |
| Serious incident reporting | EU AI Act Article 73 | Incident report with the date of awareness, root cause note, corrective action | Incident manager | Not set by Article 18; keep with the incident file |
| Use in line with instructions | EU AI Act Article 26(1) | Deployment note showing configuration matches the instructions for use | Deployer’s system owner | While the system is in use |
| Fundamental rights impact assessment | EU AI Act Article 27 | Completed FRIA, dated before first use | Deployer’s compliance lead | Not specified in the Act |
| AI literacy | EU AI Act Article 4 | Training records and role-based competence records | People or L&D lead | Not specified in the Act |
| Management system evidence | ISO/IEC 42001 | Statement of Applicability, internal audit reports, management review minutes, corrective action records | Management system owner | Certification cycle, typically three years |
| Documented risk tracking | NIST AI RMF GOVERN 1.4, MEASURE 2.1, MEASURE 3.1 | Transparent policies and procedures, TEVV documentation, tracked risk register | AI governance manager | Set by your own policy |
Two columns start most internal arguments: the owner and the retention. Both stay undecided in most programmes until an auditor asks, and both are settled faster when the register that holds them is the same one used to approve systems. The crosswalk that shows which single artefact satisfies clauses in more than one instrument is set out in where these three instruments overlap.
Article 18 sets one clock; Article 19 sets another
Read Article 18(1) closely and it names five categories only: the technical documentation under Article 11, the quality management system documentation under Article 17, documentation on changes approved by notified bodies, decisions issued by notified bodies, and the EU declaration of conformity under Article 47. Each has to stay available to national authorities “for a period ending 10 years after the high-risk AI system has been placed on the market or put into service”.
Log evidence runs on a different clock. Article 19(1) requires providers to keep automatically generated logs “for a period appropriate to the intended purpose of the high-risk AI system, of at least six months”, and Article 26(6) repeats that floor for deployers.
Evidence that falls between those two clocks is where programmes fail quietly. A risk assessment, a bias test result or a monitoring report is not named in Article 18, so its retention comes from whichever file it belongs to, usually Annex IV. An engineering team that rotates logs every 30 days breaks Article 19 without reading it, and nobody notices until an incident four months old needs reconstructing. Instrument-by-instrument periods, including the ones outside the EU, sit in how long each record has to survive.
Providers prove the build; deployers prove the use
Buying an AI system changes which records you hold, not whether you hold any.
Providers own the technical file, the conformity work and the declaration. Deployers own proof that the system ran the way the provider specified. Article 26(1) requires deployers to use the system in line with the instructions for use, which turns those instructions into a record you keep and follow. Article 26(5) requires monitoring and prompt notification of the provider when a risk or serious incident appears, producing a monitoring note and a notification trail. Article 26(6) sets the six-month log floor. Article 27 adds a fundamental rights impact assessment for public bodies and for private operators running creditworthiness assessment or life and health insurance pricing, completed before first use.
Procurement decides how hard this gets. Ask for the instructions for use, the stated intended purpose and the logging interface before signing, because the duty to monitor starts when the system goes live, not when a vendor replies to an email.
Evidence fails on ownership and timing before it fails on content
Evidence separates from draft on three attributes: attributable, dated, and produced in the course of the work rather than for the audit. A risk assessment with no author, no date and no approval is a draft, whatever its file name says.
An approval that holds up carries five fields: the system and its version, the decision taken, the role that took it, the date, and the artefacts reviewed before the decision. Name roles rather than individuals, since “approved by the AI governance manager on 14 March 2026” survives a resignation and “approved by Priya” does not. The opening questions an auditor uses, and the record that answers each, are collected in the questions an auditor opens with.
Evidence timing is the other failure mode, and it shows. An auditor who receives four documents created in the same week, describing decisions taken eighteen months earlier, has learned something about the programme rather than about the system.
Public enforcement data shows how routinely the paperwork is simply absent. In an audit report dated 2 December 2025, the New York State Comptroller examined enforcement of Local Law 144 in New York City and identified 17 potential instances of non-compliance among the 32 employers examined, against two complaints filed in two years. The Local Law 144 duty is narrow, a published bias audit summary and a candidate notice, and it was still unmet across a sizeable share of the sample.
Serious incidents come with a clock measured in days
Most evidence deadlines are measured in years. Incident reporting is measured in days, which is why it needs its own runbook rather than a paragraph in a policy.
Article 73(2) gives providers 15 days from becoming aware of a serious incident to report it to the market surveillance authority of the Member State where it occurred. Article 73(3) cuts that to two days for a widespread infringement or certain serious incidents. Article 73(4) sets 10 days where a person has died.
Both windows depend on two things existing beforehand: a log that shows what the system did, and a named person who owns the decision to report. Both are records, and both are usually missing on the day they are needed. The record that proves a human was in the loop, rather than merely available, is covered in proving a human reviewed the decision.
Machine-readable assurance is where this is heading
Evidence is moving from documents to data. A paper submitted on 15 April 2026, Making AI Compliance Evidence Machine-Readable, states the gap plainly. Frameworks such as the EU AI Act, ISO/IEC 42001 and the NIST AI RMF “specify what to assure but provide no executable format for how”, its authors write. Their proposal adapts OSCAL, the NIST format already used for FedRAMP, adds 16 property extensions for lifecycle and risk traceability, and generates assurance evidence as a byproduct of model training. The team tested it on two Annex III systems: a credit scoring model and a medical imaging segmentation model.
For a compliance lead the format matters less than the direction. Records that can only be assembled by hand, out of slides and email threads, cost more every year the portfolio grows.
Where records live decides whether you can produce them
Evidence in a spreadsheet survives about three systems and one owner. Past that, the failure is predictable: the register goes stale, the approval sits in email, the test result sits in a notebook, and nobody can say which model version produced the decision a customer is disputing.
Govern365 is built around that problem. The AI System Registry holds each system with its risk tier, owner and version. The Audit Evidence Manager stores each artefact against the obligation it answers, with the date and approver attached. Governance Workflows route a system through review so the approval is generated inside the process instead of written up afterwards, and Continuous Monitoring keeps the post-market record accumulating rather than being reconstructed later. All four are features of one platform at govern365.ai, built on the view that a record produced late costs more and proves less. How a single system moves from intake to approved, with the artefact attached at each step, is shown on where these approvals are tracked.
A 30-day evidence baseline
Programmes that pass their first review rarely start with documentation. Baselines start with a list and a sample.
- List the systems. Name, owner, purpose, risk tier, and whether you are provider or deployer for each. Documentation on a system nobody has listed cannot be found in an audit.
- Pick one high-risk system and assemble its file. Annex IV sections, risk assessment, test results, instructions for use. The gaps in that single file predict the gaps in the rest.
- Check the log clock. Confirm retention meets the six-month floor in Articles 19 and 26(6) before anyone needs a log from four months ago.
- Assign owners by role. One named role per record type, written into the register rather than agreed in a meeting.
- Wire approvals into the workflow. If the approval happens in a ticket, the ticket is the record. Contemporaneous capture removes the reconstruction problem permanently.
- Write the incident runbook. Who decides, who reports, and against which of the 15, 10 or two-day windows in Article 73.
Evidence lists go deeper than one page. The full pre-audit list, with the instrument behind each item, is set out in the records to have ready before an audit, and the documentation set underneath it is described in the records that make a system defensible.
Frequently asked questions
What counts as AI compliance evidence?
Audit evidence is any record showing an obligation was met: technical documentation, event logs, risk assessments, approvals, test results, training records, monitoring data and incident reports. Each needs a date, an author or approver, and a version. Policies describe the intent behind those records, so they support the evidence without being it.
How long do we have to keep AI governance records?
Article 18(1) of the EU AI Act sets 10 years after market placement for five categories: technical documentation, quality management system documentation, notified body change approvals, notified body decisions and the EU declaration of conformity. Automatically generated logs run on a separate floor of at least six months under Articles 19 and 26(6).
Who owns AI compliance evidence?
Evidence ownership is assigned per record, not per programme. Technical documentation usually sits with the product or engineering lead, risk assessments with the AI governance manager, logs and monitoring with the system owner, and incident reports with incident response. An auditor tests whether the named owner can produce the record on request.
Does ISO 42001 certification prove EU AI Act compliance?
No. Certification shows an AI management system meets ISO/IEC 42001, which covers much of the governance machinery an auditor wants to see. Certification does not cover EU-specific duties such as Annex IV technical documentation, registration in the EU database under Article 49, or serious incident reporting under Article 73.
Do we need evidence before the December 2027 high-risk deadline?
Yes, for anything launching before then. Article 11 requires technical documentation to exist before a system is placed on the market or put into service, and logging has to be designed in rather than switched on later. Which dates moved and which did not is set out in the dates the omnibus moved.
What does an auditor ask for first?
Usually the AI system inventory, then a sample. Pick one system and show the risk assessment, the approval, the test results, the instructions issued, the logs from a named date and the last review. Evidence that can be produced for one system within an hour generally exists for the rest of the portfolio.
