Ask a governance committee whether it has oversight of AI systems and the answer is usually yes, since a policy exists and the committee meets. Ask instead for the inventory, the named owner of one system chosen at random, the conformity file and the drift baseline recorded at release, and the answer changes. This instrument scores twenty controls across four domains out of a hundred points, with an evidence column against every row. A control counts only when the artefact can be produced inside the meeting.
Add the four domain subtotals and read the band. The number is a conversation starter rather than a compliance position, but the band tells you whether your next external request gets answered with a file or with a meeting.
You can answer an auditor, a regulator or an enterprise customer with a file rather than a meeting. The work from here is keeping the drills running and widening coverage to the systems you bought rather than the ones you built, since AI features arriving inside purchased software are the entries that go missing from a register.
The programme is real and the evidence is partial, which is the most common place to land. Pick the failed rows that are documents and drills rather than engineering work, and close them before the next release adds systems to the register. A score that moves because the estate shrank is not an improvement.
You are governing on intent, whatever the policy says, and you sit with the 63% who had no working AI governance in 2025. Start with the register and the named owners. They are the cheapest rows on the scorecard, they carry 14 of the 30 points in Domain A, and every other domain depends on them.
A2 is worth 7 points and costs a column plus three confirmation emails. Mail three owners at random and count the replies that accept the accountability, since a team alias cannot be asked anything.
C2 is worth 5 points and the review tool usually produces the number already. Examine the reviewers sitting at zero overrides rather than praising them, since a reviewer who never disagrees is a rubber stamp.
B4 is worth 4 points and takes half a day. Someone outside the team names a system with no notice and asks for its complete file within five working days, and the dated drill record is the artefact.
D4 is worth 4 points. Name the regulator, the deadline and the person who files, then run one tabletop inside twelve months, since a reporting clock is a poor place to discover nobody owns the filing.
Inventory and ownership takes 30, documentation and evidence 25, monitoring and incident response 25, and oversight and decision authority 20. Each domain holds five controls weighted between 2 and 7 points. The weights follow what gets requested first and what takes longest to reconstruct, not what costs most to fix.
No, and that rule is what makes the score useful. Full weight or nothing, judged on whether the named artefact can be produced in the room. Half marks are how a scorecard drifts back into a maturity survey, where everybody lands comfortably in the middle and nothing gets scheduled.
It is an ordinary one, which is the point of publishing it. That insurer had a committee, a policy and eleven production systems, and still sat in the bottom band. Four fixes costing no engineering moved it to 68 inside a quarter, so the distance from ordinary to defensible is short.
No. Compliance is assessed against your systems by people who did not fill in your scorecard. What the number gives a board is a shape to argue about and a programme a place to begin. The only real test is whether the evidence exists when somebody asks, unannounced, about a system you did not choose.
Twice a year, and once more in the week before your first external request for a conformity file. Record the date and who was in the room, since the next score is compared against this one, and a score that improved after a different group answered tells you nothing.
This one measures where you stand. AI Governance Under Regulation sets out the duties and the dates, AI Governance: An Engineering Reference gives the control catalogue and the artefact that proves each one, and AI Governance: What Buyers Should Ask turns the same evidence into vendor questions. Score first, then pick.
Head of AI, CISO or head of risk, with the people who hold the artefacts in the room rather than on a distribution list. One hour, the evidence column open, and somebody willing to say no. Scoring it alone at a desk produces a number that nobody else will recognise.
Drop your details and we'll send Benchmarking AI Governance straight to your inbox - no spam, unsubscribe anytime.
Bring your AI inventory, or the absence of one, to a working session with our engineering leads. We will score it, name the artefacts that do not exist, and rank the fixes by points per pound. SECTION 7 - FAQ - 5 to 8 questions
Book a scoring session