Logiciel Solutions Contact Us
Success Stories Tech News Contact Us
whitepaper

Benchmarking AI Governance.

Ask a governance committee whether it has oversight of AI systems and the answer is usually yes, since a policy exists and the committee meets. Ask instead for the inventory, the named owner of one system chosen at random, the conformity file and the drift baseline recorded at release, and the answer changes. This instrument scores twenty controls across four domains out of a hundred points, with an evidence column against every row. A control counts only when the artefact can be produced inside the meeting.

In depth

Self-Assessment Measures Intent, And Intent Is Not What Gets Requested.

01

Award the full weight or nothing, with no partial credit.

A control scores only when the named artefact exists and somebody in the room can produce it while the meeting is still running, so a policy in draft, a roadmap item or a supplier assurance scores zero. The evidence column exists to make optimism impossible to score. Run it with the people who would actually be asked, not with the people who wrote the policy.

In shortnot with the people who wrote the policy
02

The weights follow what gets requested first.

Inventory and ownership carries 30 points, documentation and evidence 25, monitoring and incident response 25, and oversight and decision authority 20, which reflects what an auditor asks for at the start and what takes longest to reconstruct, rather than what costs most to fix. Domains A and B are where almost every score is lost. Both are decided in the first fortnight of a build and remediated expensively at the end.

In shortBoth are decided in the first fortnight of a build a…
03

The shape of a score matters more than the number.

A European insurer with an active AI committee, a published policy and eleven production systems scored 48 out of 100, losing fourteen points it had expected to keep, and the register turned out to be reasonably complete with no entry naming a person. Four fixes worth 20 points between them needed a column, three confirmation emails, a report the review tool already produced, a half-day drill and one tabletop. None of the four needed engineering or a budget approval.

In shortNone of the four needed engineering or a budget approval
The detail

What Your Total Actually Tells You.

Add the four domain subtotals and read the band. The number is a conversation starter rather than a compliance position, but the band tells you whether your next external request gets answered with a file or with a meeting.

Zone · 01

Eighty and above

You can answer an auditor, a regulator or an enterprise customer with a file rather than a meeting. The work from here is keeping the drills running and widening coverage to the systems you bought rather than the ones you built, since AI features arriving inside purchased software are the entries that go missing from a register.

Zone · 02

Fifty-five to seventy-nine

The programme is real and the evidence is partial, which is the most common place to land. Pick the failed rows that are documents and drills rather than engineering work, and close them before the next release adds systems to the register. A score that moves because the estate shrank is not an improvement.

Zone · 03

Under fifty-five

You are governing on intent, whatever the policy says, and you sit with the 63% who had no working AI governance in 2025. Start with the register and the named owners. They are the cheapest rows on the scorecard, they carry 14 of the 30 points in Domain A, and every other domain depends on them.

By the numbers

The figures that make it a board-level conversation.

48 / 100
the score a European insurer reached, with a committee, a policy and eleven live systems
63%
had no AI governance policy of any kind the year before
+20
points the four cheapest fixes were worth, inside a single quarter
Inside the report

What you'll take away.

01

Step 1 - Put a human name against every register entry

A2 is worth 7 points and costs a column plus three confirmation emails. Mail three owners at random and count the replies that accept the accountability, since a team alias cannot be asked anything.

02

Step 2 - Report the override rate you already collect

C2 is worth 5 points and the review tool usually produces the number already. Examine the reviewers sitting at zero overrides rather than praising them, since a reviewer who never disagrees is a rubber stamp.

03

Step 3 - Run a cold file drill on one random system

B4 is worth 4 points and takes half a day. Someone outside the team names a system with no notice and asks for its complete file within five working days, and the dated drill record is the artefact.

04

Step 4 - Write the incident runbook and rehearse it once

D4 is worth 4 points. Name the regulator, the deadline and the person who files, then run one tabletop inside twelve months, since a reporting clock is a poor place to discover nobody owns the filing.

Questions

Frequently asked.

How is the hundred points split across the four domains?

Inventory and ownership takes 30, documentation and evidence 25, monitoring and incident response 25, and oversight and decision authority 20. Each domain holds five controls weighted between 2 and 7 points. The weights follow what gets requested first and what takes longest to reconstruct, not what costs most to fix.

Can we award half marks where a control is partly in place?

No, and that rule is what makes the score useful. Full weight or nothing, judged on whether the named artefact can be produced in the room. Half marks are how a scorecard drifts back into a maturity survey, where everybody lands comfortably in the middle and nothing gets scheduled.

Is 48 out of 100 a bad score?

It is an ordinary one, which is the point of publishing it. That insurer had a committee, a policy and eleven production systems, and still sat in the bottom band. Four fixes costing no engineering moved it to 68 inside a quarter, so the distance from ordinary to defensible is short.

Does a good score mean we are compliant?

No. Compliance is assessed against your systems by people who did not fill in your scorecard. What the number gives a board is a shape to argue about and a programme a place to begin. The only real test is whether the evidence exists when somebody asks, unannounced, about a system you did not choose.

How often should we score it again?

Twice a year, and once more in the week before your first external request for a conformity file. Record the date and who was in the room, since the next score is compared against this one, and a score that improved after a different group answered tells you nothing.

How does this relate to your other AI governance papers?

This one measures where you stand. AI Governance Under Regulation sets out the duties and the dates, AI Governance: An Engineering Reference gives the control catalogue and the artefact that proves each one, and AI Governance: What Buyers Should Ask turns the same evidence into vendor questions. Score first, then pick.

Who should fill this in?

Head of AI, CISO or head of risk, with the people who hold the artefacts in the room rather than on a distribution list. One hour, the evidence column open, and somebody willing to say no. Scoring it alone at a desk produces a number that nobody else will recognise.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Benchmarking AI Governance straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Score the twenty controls with us, against your own systems.

Bring your AI inventory, or the absence of one, to a working session with our engineering leads. We will score it, name the artefacts that do not exist, and rank the fixes by points per pound. SECTION 7 - FAQ - 5 to 8 questions

Book a scoring session