Field workbook · ISO/IEC 42001 × ISO 19011:2026 × ISO/IEC 42005
Your First
42001 Audit
An end-to-end workbook for the auditor who has been handed an AI management system and a date. Criteria, method, question banks, evidence lists, and how to write a finding that holds.
- Criteria come from
- ISO/IEC 42001:2023
- Method comes from
- ISO 19011:2026
- Subject matter
- ISO/IEC 42005:2025
- Audit type
- First or second party
- Checkboxes
- Save in this browser
Your progress
0 / 0 steps
This is written for a first-party (internal) or second-party (supplier) audit of an AI management system. Third-party certification audits are governed by ISO/IEC 17021-1 and have requirements this workbook does not cover.
Clause numbers, control references and definitions come from the publicly viewable sections of the three standards on the ISO Online Browsing Platform. The requirement and guidance text of all three is paywalled and none of it is reproduced here — every question, evidence list and technique below is an implementation reading written to the clause titles and to established audit practice. Buy the standards before you audit against them; you cannot audit against criteria you have not read.
Part 0
Orientation
Four things to have straight before you do anything else. Half of what goes wrong in a first audit is decided here, silently, before anyone opens a document.
0.1What this combines
Three standards do three different jobs in a 42001 audit, and confusing them is the fastest way to produce a report nobody can act on.
| Standard | Its job here | You use it to… |
|---|---|---|
| ISO/IEC 42001:2023 Criteria |
The requirements the auditee must meet — clauses 4 to 10, plus every Annex A control they marked applicable | Decide what conformity means. Every finding you raise must cite a specific requirement from here, from their own documents, or from law. |
| ISO 19011:2026 Method |
How to run the audit — principles, programme, the six stages of an audit, competence | Decide how you do it: planning, sampling, interviewing, evidence, findings, reporting, follow-up. |
| ISO/IEC 42005:2025 Subject matter |
What a good AI system impact assessment contains | Judge the quality of what you find at 42001 clauses 6.1.4 and 8.4 and controls A.5.2 to A.5.5. Without it you can confirm an assessment exists but not whether it is any good. |
There is a fourth input that is easy to forget and is often where the sharpest findings come from: the auditee's own documents. Their AI policy, their procedures, their contractual commitments to customers. These are criteria too, under ISO 19011 clause 3.8, and an organisation failing its own stated policy is a cleaner finding than one failing a standard's general wording.
0.2The chain you never break
Five defined terms in a strict order. Every weak audit you will ever read broke this chain somewhere, usually at the first link.
| Order | Term | Test |
|---|---|---|
| 1 | Audit criteria | Can you point to the exact clause, control, policy section or legal provision? If not, you have an opinion. |
| 2 | Audit evidence | Is it verifiable — could a second auditor go and look at the same thing? Records, statements of fact, information. Note the reference: document ID, version, date, system, who said it. |
| 3 | Audit finding | Evidence evaluated against the criteria. Conformity, nonconformity — or an opportunity for improvement, or good practice worth recording. |
| 4 | Audit conclusion | One per audit, reached after considering the objectives and all findings. Not the sum of the bad news. |
| 5 | Follow-up | Not closure — effectiveness. Did the action actually work? |
Walking in with a conclusion and collecting evidence that supports it. It happens most to auditors who know the organisation well, because they already suspect where the problems are. Suspicion is a legitimate input to sampling — go and look there first. It is not a legitimate input to findings. If the evidence does not support what you expected, the finding is that the evidence does not support it.
0.3Are you fit to audit this?
ISO 19011 clause 7 asks this formally. Answer it honestly now, because the alternative is discovering the gap in an interview when someone explains model validation to you.
Not the technical one — the impact one. Auditors with an information security background arrive fluent in controls and evidence and audit an AIMS beautifully as though it were an ISMS, never once asking what the system does to the people it makes decisions about. That question is the whole difference between the two standards, and if nobody on the team is equipped to ask it, the audit will pass a system that has never been assessed for harm.
0.4Words to use precisely
Auditees mirror your vocabulary. Use these accurately from the opening meeting and the whole audit gets easier.
| Term | Means | Not |
|---|---|---|
| Nonconformity | Non-fulfilment of a requirement | Something you would have done differently |
| Correction | Fixing the instance | Corrective action |
| Corrective action | Eliminating the cause so it does not recur | Retraining the person who made the mistake |
| Opportunity for improvement | Conforms, but could be better | A soft nonconformity you did not want to argue about |
| Compliance / non-compliance | Used when the criteria are legal or regulatory | Interchangeable with conformity |
| AI risk assessment | What could happen to the organisation — 42001 clauses 6.1.2, 8.2 | The same exercise as impact assessment |
| AI system impact assessment | What could happen to individuals, groups and societies — 42001 clauses 6.1.4, 8.4 | A section of the risk register |
| Statement of Applicability | All necessary controls plus justification for every inclusion and exclusion | A list of Annex A with ticks |
| Technical expert | Advises the team; does not audit | An extra auditor |
Part 1
Before the audit
Seven steps, over the two to three weeks before you walk in. Preparation is where audit quality is actually determined — a well-prepared auditor with two days finds more than an unprepared one with five.
Step 1Confirm the assignment
Three things must be written down and agreed before anything else happens: objectives, scope, criteria. ISO 19011 clause 5.5.2. Skipping this is why audits end in arguments about whether something was in scope.
Objectives — what this audit is to establish
Write them so they could fail. Typical objectives for an AIMS audit:
- Determine conformity of the AI management system with ISO/IEC 42001 clauses 4–10 and applicable Annex A controls
- Evaluate the effectiveness of the system in achieving stated AI objectives
- Identify opportunities for improvement
- Confirm readiness for a stage 1 certification audit (only if that is genuinely the purpose — it changes the depth)
Scope — extent and boundaries
Name all of it: organisational units, functions, processes, which AI systems, physical and virtual locations (19011 clause 3.6 now expects both), and the time period covered.
Start from the AIMS scope statement (42001 clause 4.3), then decide whether your audit covers all of it or a slice. If a slice, say which — and say it in the report too.
Criteria — the exact requirements
List them explicitly. For a full AIMS audit that is: ISO/IEC 42001:2023 clauses 4–10; every Annex A control marked applicable in the current SoA (cite its version); the organisation's AI policy and AIMS procedures; and applicable legal and regulatory requirements.
Step 2Build the criteria pack
One folder. Everything you will cite a finding against, and nothing else. Assemble it before the desk review so you are reading documents against something.
Each applicable control in the SoA is a criterion with an implementation claim already attached — the auditee has written down what they do and where the evidence is. Every one of those claims is a thread you can pull. Planning an AIMS audit is largely the exercise of deciding which threads, and in what order.
Step 3Request documents in advance
Send this list a fortnight out. What arrives — and what does not, and how long it takes — is already evidence.
| Clause | Ask for | What its absence tells you |
|---|---|---|
| 4.1–4.2 | Context analysis; interested parties register | Often exists as a one-off document from the implementation project |
| 4.3 | AIMS scope statement | Must exist — it is required documented information |
| 4.4 | AI system inventory, with the organisation's role for each system | The single most diagnostic document. No inventory means no governed portfolio. |
| 5.2 | AI policy, with approval and communication evidence | Required |
| 5.3 | Roles, responsibilities and authorities | Watch for roles assigned to job titles that are vacant |
| 6.1.2 | AI risk assessment process and criteria; the risk register | Required |
| 6.1.3 | Risk treatment plan; residual risk acceptance records | Acceptance records are frequently missing |
| 6.1.3 | Statement of Applicability | Required. If it is not versioned, that is a finding waiting. |
| 6.1.4 | AI system impact assessment process | Required |
| 8.4 | Completed impact assessments for in-scope systems | A defined process with no completed assessments is the classic AIMS gap |
| 6.2 | AI objectives and plans, with latest measurement results | — |
| 7.2 | Competence matrix and training records | — |
| 8.1 / 6.3 | Change control procedure and recent AI change records | Ask specifically for model version changes and retraining events |
| 9.1 | Monitoring plan and results, including AI system-level metrics | Process metrics without system metrics is a pattern worth noting |
| 9.2 | Internal audit programme and previous reports | Required |
| 9.3 | Management review minutes | Required — check every input in 9.3.2 is covered |
| 10.2 | Nonconformity and corrective action log | An empty log after twelve months is itself a finding |
| A.10 | Supplier agreements for AI systems and models; responsibility allocations | — |
| A.8.4 | AI incident records | — |
Step 4Desk review
ISO 19011 clause 6.3.1. Read before you arrive, and read to build questions rather than to form conclusions.
Read in this order
- AIMS scope statement — what is in, what is out, and does the exclusion make sense
- AI system inventory — how many systems, what roles, which look highest-exposure
- Statement of Applicability — which controls apply, what justifications say, what status is claimed
- Risk register and impact assessments — are these two distinct exercises or one wearing two labels?
- Management review minutes and internal audit reports — is the loop running?
- Nonconformity log — what kind of problems does this organisation find in itself?
Red flags to note now and test later
| What you see | What to test on the day |
|---|---|
| SoA marks every control applicable | Ask them to walk you through three exclusion decisions they considered. If there were none, the SoA was not reasoned. |
| SoA justifications say “best practice” or “required by standard” | Pick a row and ask which risk or impact finding drove it. Inclusion should trace back. |
| Impact assessment findings read like organisational risks — reputational, legal, regulatory | Ask who bears each harm. If the answer is always “us”, clause 6.1.4 has been collapsed into 6.1.2. |
| Impact assessment process defined, no completed assessments | Straightforward: clause 8.4 requires them to be performed and retained. |
| Assessments do not name a model version | Compare against what is actually running. If it has changed, the assessment describes a system that no longer exists. |
| AI inventory has fewer entries than you would expect from the size of the organisation | Ask about embedded AI in SaaS tools and vendor features switched on by default. |
| Nonconformity log is empty | Not a clean system. A system that is not looking. |
| Corrective actions all read “staff reminded” or “training provided” | Root cause analysis is not happening. Test one recurrence. |
| Management review minutes are short and decision-free | Check every clause 9.3.2 input is present and that decisions with owners came out. |
| Every AI objective was met | Objectives that never miss are not measuring anything. Ask how the target was set. |
| Human oversight described everywhere as a safeguard | Ask for the override rate. Nominal oversight is the most under-detected weakness in an AIMS. |
Step 5Choose the sample
You cannot audit everything, and 19011's risk-based principle (clause 4.8) says you should not try. Sampling is where a new auditor most needs a deliberate strategy rather than a checklist walked top to bottom.
Sample on four axes
| Axis | Pick | Because |
|---|---|---|
| By exposure | The AI system with the highest impact rating in their own assessments | If governance holds anywhere it should hold here. If it does not hold here, nothing else matters. |
| By change | A system changed in the last quarter — retrained, new data source, new market, oversight altered | Change is where assessments go stale and clause 6.3 gets tested. |
| By newness | The most recently deployed system | Tests whether the process gates work under delivery pressure, or only in retrospect. |
| By blind spot | A third-party or embedded AI feature nobody thinks of as “their” AI | Deployer obligations are the most commonly missed part of the standard. |
Fair presentation (19011 clause 4.3) means the report states the sample. “Three of eleven AI systems were examined; the remainder were not” is honest and preserves the value of a clean conclusion. A report that reads as though everything was covered when four systems were is the kind of thing that damages an audit function years later.
Step 6Choose the method
ISO 19011:2026 clause 5.5.3, with remote auditing now a defined method (3.4) drawn from ISO/IEC TS 17012. For an AIMS this matters more than usual, because much of what you are auditing has no physical location.
Remote works well for
- Documented information review
- Sampling records and system data live — often better than on-site
- Model registries, pipelines, logs, access controls, monitoring dashboards
- Interviews with dispersed data science and engineering teams
- Auditing a virtual location — work performed in an online environment with no site
Remote works badly for
- Observing an operator actually using the AI system in the real setting
- Seeing whether human oversight is real — you need to watch the queue at a busy hour
- Following an unplanned thread; a screen-share shows what was prepared
- The unprompted aside, the hesitation, the corridor remark
- Culture and tone, which is much of what tells you where to look next
A hybrid design is usually right for an AIMS: remote for the document and system evidence, on-site for the human-oversight and operational reality. Where something cannot be verified remotely, plan for it — and if it ends up unverified, the report says so.
Step 7Write the audit plan
ISO 19011 clause 6.3.2. Send it to the auditee in advance. A plan they have seen is a plan they can staff.
Put leadership early — clause 5 sets the tone and tells you how seriously to read everything else. Put the system trace (Step 13) in the middle, when you know enough to follow threads and still have time to chase what it opens up. Leave the last session unallocated. You will need it, and an auditor with no slack stops pulling threads at exactly the point they get interesting.
Part 2
During the audit
Opening meeting to closing meeting. The question banks in this part are the working core of the workbook — take them in, but do not read from them.
Step 8Opening meeting
Fifteen to thirty minutes. Its job is to remove every ambiguity that could otherwise surface as friction later.
Say plainly that you are here to find out whether the system works, that findings are about the system rather than about people, and that you would rather be told “we do not do that” than shown something assembled this morning. You will not always be believed, but the people who do believe you will give you a far better audit — and the ones who assemble something anyway are usually detectable by the version date.
Step 9Interview technique
The skill that separates a competent auditor from a checklist-reader. Four moves, in order, every time.
Open — let them describe it
“Walk me through what happens when a new AI system is proposed.” Not “do you have a process for assessing new AI systems?” — a closed question invites a yes and teaches you nothing. The description tells you what they actually do, and the gap between it and the procedure is often the finding.
Narrow — get to a specific instance
“Take the last one that went through. Which system was it?” General descriptions are unfalsifiable. A named instance can be checked.
Ask for the evidence
“Can you show me the assessment for that one?” Then look at it properly, in front of them, and note the reference: document ID, version, date, author, approver.
Verify — is it what it claims to be?
Does the version match what is running? Is the date before deployment or after? Is the approver the person the procedure names? Did the mitigations it promised actually happen? This is the step new auditors skip, and it is the step that finds things.
Handling the four answers you will get
| They say | You do |
|---|---|
| “We do that, it's just not written down.” | Ask them to show you an instance anyway. Undocumented practice may still produce records. If it produces nothing verifiable, the gap is against whichever clause requires documented information — say which. |
| “That's handled by the vendor.” | Ask to see the contract or the responsibility allocation (A.10.2). Responsibility can be transferred; accountability for having transferred it cannot. |
| “That's not applicable to us.” | Check the SoA. If it is marked applicable, this is a divergence between the SoA and practice — itself a finding. If it is excluded, test whether the justification holds. |
| A long technical explanation you cannot evaluate | Say so, and ask for it in terms of the requirement: “Help me connect that to control A.6.2.4 — what evidence shows the validation was performed and reviewed?” Never nod through something you did not follow. That is how weak evidence gets accepted. |
The AIMS manager can describe the system as designed. The data scientist, the operator, the person who actually reviews the model's output can tell you what happens on a busy Tuesday. When those two accounts differ, the difference is the audit.
Step 10Clause question banks
Clauses 4 to 10, with what to ask, who to ask, what evidence settles it, and the weak answer to watch for. Take these in as prompts — an auditor reading from a script finds only what the script anticipated.
Clause 4 — Context
Context and interested parties
AIMS manager- What internal and external issues affect your ability to govern AI here? When did you last revisit them?
- Who are your interested parties, and what does each require of you?
- Are the people your AI systems make decisions about on that list? What are their requirements, and how did you find out?
Evidence that settles itContext analysis with a recent review date; interested parties register whose entries trace into objectives, risks or controls.
Weak answerA PESTLE grid produced during implementation and untouched since — and a parties register listing only customers, regulators and staff, with no one affected by the systems' outputs.
AIMS scope
AIMS manager, top management- What is in scope and what is excluded? Why?
- Which AI systems in the organisation are outside the AIMS, and how would a customer reading the scope statement know that?
Evidence that settles itDocumented scope statement, and an inventory that lets you check which systems fall inside it.
Weak answerA scope so narrow that most of the organisation's AI activity sits outside it, with no acknowledgement of the gap.
The system itself
AIMS manager- Show me your AI system inventory. How is it kept current?
- How does a new AI system get onto it? What if a team enables an AI feature in a tool they already have?
- How do the parts of the AIMS interact — how does a risk become a control, how does an incident become a corrective action?
Evidence that settles itA maintained inventory with roles per system, plus a process map showing the parts connecting.
Weak answerA spreadsheet last updated at implementation, and no route by which an embedded AI feature would ever appear on it.
Clause 5 — Leadership
Leadership and commitment
Top management — interview them directly- What are the AI objectives, and how do they relate to strategic direction?
- Tell me about a decision where the AIMS position and the commercial position differed. What happened?
- What resources have you provided for AI governance in the last year?
- How do you know the AIMS is working?
Evidence that settles itManagement review attendance and decisions; a resourcing decision; ideally an instance where a launch was delayed or changed on governance grounds.
Weak answer“We're fully committed” with no instance behind it — and management review minutes showing a delegate attending in their place.
AI policy
Top management; staff at random- Who approved the policy and when? When is it reviewed (A.2.4)?
- How does it relate to your other policies (A.2.3) — where do they overlap or conflict?
- Ask a developer or operator: what does the AI policy mean for your work?
- Is there anything the policy says you will not do?
Evidence that settles itApproved, versioned, communicated policy — and staff who can say something concrete about it.
Weak answerA policy that restates the standard back to itself and takes no position on anything contestable. Then nobody outside the governance team can describe it.
Roles and authorities
AIMS manager; named role-holders- Who is responsible for AIMS conformity? Who reports performance to top management?
- How does someone raise a concern about an AI system (A.3.3)? Has anyone used it?
- To a named role-holder: what are you responsible for, and what authority do you have?
Evidence that settles itDocumented assignments, and a reporting route that has carried actual traffic.
Weak answerRoles assigned to titles rather than people; a concerns channel nobody has ever used and few know exists.
Clause 6 — Planning
AI risk assessment
Risk owner, AIMS manager- Walk me through the method. What are your risk criteria?
- How do you get consistent results — would two assessors land in the same place?
- Show me the assessment for [sampled system]. Who performed it, when, and who reviewed it?
- What AI-specific risk sources did you consider? Did you use anything like Annex C as a prompt?
Evidence that settles itA documented method with criteria, and assessments demonstrably following it.
Weak answerRisks that would read identically for any IT system — availability, breach, vendor lock-in — with nothing about continuous learning, opacity, or data-derived behaviour.
Risk treatment and the SoA
AIMS manager, risk owners- Take this control marked applicable — which risk or impact finding drove its inclusion?
- Take this exclusion — on what basis is it not applicable?
- How did you compare your controls against Annex A to check nothing was overlooked?
- Who accepted the residual risk, and what authority do they have to accept it?
Evidence that settles itAn SoA that survives a random-row walk in both directions — backwards to the risk, forwards to the operating evidence — plus signed residual-risk acceptance.
Weak answer“Best practice” as a justification; every control applicable; no residual risk acceptance record at all.
Impact assessment process
Process owner- How is this process different from your risk assessment process?
- When does an impact assessment happen in the life cycle? What triggers a re-assessment?
- Who takes part? Is anyone in the room representing the people affected?
- What are your thresholds for a sensitive use? What happens when one is hit?
- Who approves, and at what level?
Evidence that settles itA documented process with triggers, roles, scales, thresholds and approval routes defined in advance. See Step 12 for testing the output.
Weak answer“It's part of our risk process.” If the two are one exercise, clause 6.1.4 is not met — and the standard would not have made them separate clauses with separate operational counterparts if it were.
Objectives and change
AIMS manager, change manager- What are your AI objectives? How are they measured, and what were the last results?
- Has an objective ever been missed? What happened?
- What counts as a change to an AI system here? Does retraining count? A vendor upgrading a model beneath you?
- Show me the last AI system change and what assessment it triggered.
Evidence that settles itMeasured objectives with a trend, and a change record showing re-assessment actually fired.
Weak answerAll objectives met every period; and a change process that treats a model retrain as a routine deployment.
Clause 7 — Support
Resources, competence, awareness
AIMS manager; staff at random- How do you document data, tooling, compute and human resources for AI (A.4.2–A.4.6)?
- What competence does an AI role need here? Who determines that?
- Who in the organisation is competent to judge fairness? To judge what happens to the affected person?
- Show me a competence gap you identified and what you did about it.
- To a random staff member: what does the AI policy require of you, and what happens if it is not followed?
Evidence that settles itCompetence matrix mapped to roles; training records; a closed gap.
Weak answerCompetence defined entirely in technical terms. A team of excellent engineers can be collectively incompetent for this clause.
Communication and documented information
AIMS manager- What do you tell users and affected parties about your AI systems (A.8.2, A.8.5)?
- How would you communicate an AI incident, and to whom (A.8.4)?
- How is AIMS documentation version-controlled, approved and retained?
Evidence that settles itVersion control and approval on the documents you have already been reading; a retention schedule that is followed.
Weak answerKey procedures in an uncontrolled shared drive with no version or approval. Easy to find, impossible to argue about.
Clause 8 — Operation
Doing it
Operational owners, engineering- Show me operational control for [sampled system] — the procedures actually used.
- How do you control externally provided AI — models, APIs, AI features inside SaaS?
- When were risk assessments and impact assessments last performed for this system, and what triggered them?
- Show me a case where an unintended change occurred and what you did.
Evidence that settles itDated, approved assessments retained as records; supplier controls that match the A.10 claims in the SoA.
Weak answerClause 6 documents exist, clause 8 records do not. A plan that was never resourced becomes visible exactly here.
Clauses 9 and 10 — Evaluation and improvement
Monitoring and measurement
AIMS manager, engineering- What do you monitor, by what method, how often?
- Which of those measures are about the AI systems themselves rather than about the process — drift, disaggregated performance, incident rate, human override rate?
- What happens when a threshold is breached? Show me a time it was.
Evidence that settles itDefined measures with methods, a trend rather than a snapshot, and an instance of a breach producing an action.
Weak answerProcess metrics only — assessments completed, training delivered — with nothing measuring how the deployed systems are actually behaving.
Internal audit and management review
Internal audit, top management- Show me the internal audit programme. How is frequency decided?
- How is auditor impartiality ensured, given the size of the team?
- Take the last management review — walk me through each of the clause 9.3.2 inputs.
- What decisions came out of it, and who owns them?
Evidence that settles itA risk-based programme; reports containing findings that are not all positive; minutes covering every required input with decisions and owners.
Weak answerAn audit calendar rather than a programme; a management review that reads as a briefing with no decisions recorded.
Nonconformity and corrective action
AIMS manager, quality- Show me the log. How many nonconformities in the last twelve months?
- Take this one — what was the root cause, and how did you determine it?
- Did you check whether similar nonconformities existed elsewhere?
- How did you verify the action was effective, not just complete?
- Has anything recurred?
Evidence that settles itA log with genuine root-cause analysis, similar-instance checks, and effectiveness verification.
Weak answerAn empty log — a system that is not looking; or every root cause recorded as human error with “staff reminded” as the action.
Step 11Annex A — the 38 controls
Only the ones marked applicable in their SoA are criteria. Sample rather than sweep — but sample across all nine areas, because coverage is what stops you auditing only what you find comfortable.
| Ref | Control | Evidence to ask for |
|---|---|---|
| A.2 — Policies related to AI | ||
| A.2.2 | AI policy | Approved, versioned, communicated policy |
| A.2.3 | Alignment with other organizational policies | Mapping to security, privacy, HR policies; a resolved conflict |
| A.2.4 | Review of the AI policy | Review record with date and outcome |
| A.3 — Internal organization | ||
| A.3.2 | AI roles and responsibilities | Documented assignments to named people |
| A.3.3 | Reporting of concerns | The channel, its communication, and evidence it has been used |
| A.4 — Resources for AI systems | ||
| A.4.2 | Resource documentation | The resource record itself |
| A.4.3 | Data resources | Datasets identified, owned and described |
| A.4.4 | Tooling resources | Tool inventory; approval route for new tools |
| A.4.5 | System and computing resources | Compute and environment records |
| A.4.6 | Human resources | Roles, competence, capacity |
| A.5 — Assessing impacts of AI systems · see Step 12 | ||
| A.5.2 | AI system impact assessment process | Documented process with triggers, scales and thresholds |
| A.5.3 | Documentation of AI system impact assessments | Completed assessments in a consistent template |
| A.5.4 | Assessing AI system impact on individuals or groups | Named affected groups; harms with mechanisms; ratings |
| A.5.5 | Assessing societal impacts of AI systems | Societal and environmental dimensions actually addressed |
| A.6 — AI system life cycle | ||
| A.6.1.2 | Objectives for responsible development | Stated objectives per system or programme |
| A.6.1.3 | Processes for responsible design and development | The SDLC with AI-specific gates |
| A.6.2.2 | AI system requirements and specification | Requirements including non-functional and fairness criteria |
| A.6.2.3 | Documentation of design and development | Design records for the sampled system |
| A.6.2.4 | Verification and validation | Test and evaluation results, disaggregated by segment |
| A.6.2.5 | AI system deployment | Deployment approval and release record |
| A.6.2.6 | Operation and monitoring | Live monitoring, thresholds, and what a breach triggers |
| A.6.2.7 | Technical documentation | Current technical documentation matching the running version |
| A.6.2.8 | Recording of event logs | Logs, retention, and who reviews them |
| A.7 — Data for AI systems | ||
| A.7.2 | Data for development and enhancement | What data was used, for what, under what basis |
| A.7.3 | Acquisition of data | Source, licence, lawful basis |
| A.7.4 | Quality of data | Quality measures against a stated requirement, not in the abstract |
| A.7.5 | Data provenance | Traceable lineage to source |
| A.7.6 | Data preparation | Preparation and labelling process, and its known biases |
| A.8 — Information for interested parties | ||
| A.8.2 | System documentation and information for users | What users are actually given, in a form they can use |
| A.8.3 | External reporting | What is published or reported, and to whom |
| A.8.4 | Communication of incidents | The procedure, and a real incident communicated |
| A.8.5 | Information for interested parties | Information reaching people who are not customers |
| A.9 — Use of AI systems | ||
| A.9.2 | Processes for responsible use | Operating procedures for the people who use it |
| A.9.3 | Objectives for responsible use | Stated use objectives |
| A.9.4 | Intended use of the AI system | Intended use stated, and out-of-scope use controlled |
| A.10 — Third-party and customer relationships | ||
| A.10.2 | Allocation of responsibilities | Written allocation — not an assumption about who does what |
| A.10.3 | Suppliers | Supplier assessment and contractual AI terms |
| A.10.4 | Customers | What customers are told and committed to |
A deployer who builds nothing will legitimately exclude much of A.6, and their centre of gravity is A.9, A.10 and A.5. A developer's is A.6, A.7 and A.4. If the SoA does not reflect the organisation's actual role, either the SoA was copied from a template or the organisation has not decided what it is — and both are worth a finding.
Step 12Auditing the impact assessment
The AI-specific heart of the audit, and the part a general management system auditor is least equipped for. This is where ISO/IEC 42005 becomes your reference — not as criteria, but as the basis for judging whether what you are looking at is real.
Confirming an assessment exists takes five minutes. Judging whether it is any good is the actual work. Take one completed assessment for a high-exposure system and test it against these twelve questions.
“Who gets the benefit, and who bears the harm?” When the benefit accrues to the organisation in aggregate and the severe harm falls, rarely, on an individual who will never know it happened, that asymmetry is exactly what an impact assessment exists to make visible — and exactly what an averages-based metrics dashboard reports as success. If the assessment does not surface it, ask why.
Step 13Trace one system end to end
The highest-yield ninety minutes of the audit. Take one deployed AI system and follow it through every clause. This finds what reading policies never finds.
Between the governance documentation and the deployed reality. The model that was retrained without re-assessment. The oversight step quietly dropped for throughput. The review date that passed six months ago. The mitigation that was written down and never funded. None of these are visible from the policy shelf; all of them are visible from this trace.
Step 14The evidence log
Keep it as you go. Reconstructing evidence references from memory the following week is how findings become unverifiable — and an unverifiable finding will be dismissed the moment it becomes inconvenient.
| Field | Record | Example |
|---|---|---|
| Ref | Sequential ID | E-014 |
| Date / time | When observed | 2026-09-14, 11:20 |
| Criterion | The clause, control or policy section | 42001 clause 8.4; SoA control A.5.3 |
| Source | Document ID and version, record ID, system, or person and role | AIIA-007 v1.2, approved 2025-11-03 |
| Observation | What you actually saw or were told — factual, no judgement | Assessment references model v2.1. Engineering confirmed v2.4 in production since 2026-06. |
| Verified | How you confirmed it | Model registry viewed on screen with [name], 11:24 |
| Status | Conformity / potential nonconformity / OFI / good practice / follow up | Potential NC |
Raise potential findings as they emerge (19011 clause 6.4.4), not at the closing meeting. It gives the auditee a chance to produce evidence you had not seen — which is the point, not a concession. Surprises at a closing meeting produce defensiveness and cost you the follow-up.
Debrief the team daily. Ten minutes: what did each of us find, what threads are open, what do we chase tomorrow. Patterns across auditors are usually more significant than any single observation.
Step 15Closing meeting
Present findings in the four-part structure below. Do not soften a nonconformity into an opportunity because the room is uncomfortable — that is the failure of fair presentation (19011 clause 4.3) and it is far more costly to the audit function than a difficult ten minutes.
Equally, do not prescribe the fix. The finding states the gap; the auditee determines the cause and the action. An auditor who supplies the remedy has taken ownership of a problem that is not theirs, and has made themselves un-independent for the follow-up.
Say so, explicitly, in the meeting and in the report. Access refused, evidence unavailable, a location that could not be reached remotely, a person unavailable. Fair presentation requires reporting significant obstacles and matters left unresolved. An audit that quietly concludes on things nobody actually saw is worse than one that reports the gap.
Part 3
After the audit
Findings, grading, report, follow-up. Most of the audit's value is created or destroyed here — and a late report has already lost most of it.
Step 16Write the findings
Four parts, every time. A finding that cannot survive being questioned costs the audit function credibility it will need later.
The requirement
The specific criterion, cited precisely — clause, control reference, policy section, statutory provision. Accurate enough that the auditee can find it. If it is not in the criteria, it is an opportunity, not a nonconformity.
The evidence
What you saw, with identifiers: document reference and version, record ID, date, system, who said it and in what role. Verifiable — someone else could look at the same thing.
The gap
One sentence stating how the evidence fails to meet the requirement. No hedging, no adjectives, no context that never quite says what is wrong.
The grade
Major, minor, opportunity for improvement, or good practice. Consistent across the programme so a grade means the same thing in every report.
Worked example
Weak
“Impact assessments were found to be inadequate in some cases and should be improved. There is a risk that AI systems are deployed without proper consideration of their effects.”
No criterion. No verifiable evidence. “Some cases” is unanswerable. It reads as an opinion, and it will be treated as one.
Sound
Criterion: ISO/IEC 42001:2023 clause 8.4 — impact assessments performed at planned intervals and when significant changes occur, with results retained.
Evidence: AIIA-007 v1.2 (approved 2025-11-03) for the arrears prioritisation system references model version 2.1. The model registry, viewed with [name, role] on 2026-09-14, records v2.4 in production since 2026-06-12. No impact assessment has been performed since v1.2.
Gap: A significant change to the AI system did not trigger the impact assessment required by clause 8.4; the retained assessment describes a version no longer in production.
Grade: Major — a required process has not operated, and the system currently in production is unassessed.
Step 17Grade them
ISO 19011 does not prescribe a grading scale. Your programme should define one and apply it consistently — inconsistent grading is the fastest way for an audit function to lose authority.
| Grade | Means | Typical AIMS example |
|---|---|---|
| Major | A required process is absent, has broken down, or the failure means the system cannot achieve its intended results | No impact assessment process at all; production systems unassessed; the SoA does not correspond to what is implemented |
| Minor | A single lapse in an otherwise functioning process | One assessment missing its approval signature; a document not versioned; one review date passed |
| OFI | Conforms, but could be stronger. Never a downgraded nonconformity. | Impact assessments are sound but no affected-party consultation is used; monitoring is adequate but not disaggregated |
| Good practice | Worth recording and worth copying | Human override rate monitored as a live signal with a threshold that triggers review |
Record good practice. ISO 19011 clause 3.11 allows it and almost nobody uses it. A report containing only nonconformities trains the organisation to treat audit as a threat, and an audit people defend against yields far less than one they help with.
Step 18The report
Complete, accurate, concise, clear — and on the agreed date. Findings act on decisions being made now; a report arriving weeks later arrives after them.
Step 19Follow-up
ISO 19011 clause 6.7. The stage where internal audit programmes most reliably decay: findings raised, actions logged, effectiveness never checked.
The auditee owns the correction, the corrective action and the improvement. The audit programme verifies. And it verifies effectiveness, not completion — an action marked closed that did not work is still an open finding.
Part 4
Reference
What to expect to find, what to watch in yourself, and the last check before you put your name on it.
4.1Common nonconformities in an AIMS
What audits of AI management systems actually turn up, roughly in order of how often. Use it to aim your sampling, never as a list of findings to go and confirm.
| Clause | Finding | Usual grade |
|---|---|---|
| 6.1.4 / 8.4 | Impact assessment collapsed into risk assessment — every finding is an organisational exposure | Major |
| 8.4 | Process defined, no completed assessments for production systems | Major |
| 6.3 / 8.4 | Model changed or retrained with no re-assessment; assessment names an obsolete version | Major |
| 4.4 | AI inventory incomplete — embedded and third-party AI absent | Major |
| 6.1.3 | SoA justifications do not trace to any risk or impact finding | Minor to Major |
| 6.1.3 | Residual risk acceptance not recorded, or accepted by someone without the authority | Minor |
| A.5.4 | Harms recorded without a named affected group or a mechanism | Minor |
| A.5.5 | Societal and environmental dimensions not addressed at all | Minor |
| 9.1 | Monitoring covers the process but not the AI systems — no drift, no disaggregated performance, no override rate | Minor |
| 10.2 | Root cause recorded as human error; corrective action is “staff reminded” | Minor |
| 10.2 | Nonconformity log empty after twelve months of operation | Minor |
| A.10.2 | Vendor assumed responsible for something with nothing in the contract saying so | Minor |
| 7.5 | Key AIMS documents uncontrolled — no version, no approval | Minor |
| 9.3 | Management review missing required clause 9.3.2 inputs, or producing no decisions | Minor |
| A.6.2.4 | Validation reported as a single headline accuracy figure, not disaggregated | Minor |
| 5.1 | Top management commitment on paper; delegate attends management review | Minor |
4.2Traps for new auditors
Being led
A skilled auditee will walk you through their best-prepared area at a comfortable pace until the time runs out. You set the schedule and you choose the sample. Say so, pleasantly, and then do it.
Accepting a plausible explanation
The single biggest difference between an experienced and an inexperienced auditor. Plausible explanations are exactly what a weak process produces, because the person has had to explain it many times. Ask for the evidence — every time, including when you believe them.
Nodding through what you did not follow
Especially with technical material, and especially when you are conscious of being the least technical person in the room. Say “I need to connect that to the requirement” and make them do it. Nobody has ever thought less of an auditor for asking; plenty of weak evidence has been accepted by auditors who did not.
Auditing the paperwork instead of the system
The characteristic AIMS failure. Policies read beautifully, procedures are complete, and nobody looked at a deployed model. Step 13 exists for this.
Going native
You spend three days with people who are working hard on something difficult, you come to like them, and the findings soften. Write the finding before the sympathy has a chance to work — draft it the evening you observe it.
Over-collecting
Forty pieces of evidence and no findings. Evidence exists to answer a question against a criterion. If you cannot say which criterion a document speaks to, you do not need it.
Confusing your role with the auditee's
You state the gap. They determine the cause and the fix. Prescribing the remedy feels helpful, removes their ownership, and disqualifies you from verifying it.
Treating the checklist as the audit
A checklist is a memory aid. The best findings in any audit come from a thread that the checklist did not anticipate — someone's aside, a date that does not line up, a version number that does not match. Leave room to follow them.
4.3Before you sign
Ten minutes with the draft report, on your own.
The last one. Everything else in this workbook can be learned from any management system audit. That question is the entire reason ISO/IEC 42001 exists as a separate standard rather than an annex to ISO 27001 — and an AIMS audit that never asks it has audited an information security system with different vocabulary.
4.4Sources and companion guides
- ISO/IEC 42001:2023, ISO 19011:2026 and ISO/IEC 42005:2025 on the ISO Online Browsing Platform — forewords, introductions, scopes, terms and bibliographies are publicly viewable; the requirement and guidance text is not.
- Deep reference guides in this series: ISO/IEC 42001 — the AI management system · ISO 19011:2026 — auditing management systems · ISO/IEC 42005 — AI system impact assessment.
- For third-party certification audits, requirements are in ISO/IEC 17021-1. For remote auditing method detail, ISO/IEC TS 17012:2024.
ISO/IEC 42001:2023, ISO 19011:2026 and ISO/IEC 42005:2025 are copyrighted works of ISO and IEC. Nothing here reproduces their requirement or guidance text; clause and control titles are cited as references, and publicly viewable scope and definition material is paraphrased. This is an independent working aid, not a substitute for the standards, and not endorsed by ISO or IEC. You cannot audit against criteria you have not read — obtain licensed copies before conducting an audit.