Most AI governance assessments are answered from memory in a leadership meeting. Nobody sets an evidence bar, so intent, practice and enforcement all collapse into the same number, and eight domains get averaged into one figure that goes upward as a trend. Then an enterprise customer asks not for your score but for the artifact behind three statements they pick themselves, and the room cannot produce any of the three. The question was never "how mature do we feel". It is "what could we show somebody", and that has a measurable answer.
Three design choices do the work. Each one is the reason a score produced this way survives somebody checking it.
Zero is absent and unowned. One is intent with no owner and no date. Two is informal, happening because one person does it and stopping when they get busy. Three is defined: written, assigned to a role, followed in the normal case. Four is verified, enforced by a gate in CI, an access control or a review with a record. The step from two to three is documentation. The step from three to four is what survives an audit.
not the average
A total of 104 built from eight scores of 13 is a different company from a 104 built from five scores of 18 and three of 4.7, and the second is in more trouble. Reviews, regulators and incidents all probe the weakest domain rather than the mean. Any single domain below 10 out of 20 is the headline finding regardless of the total.
Consumer chat tools, a marketing team's unreviewed vendor, an engineer's personal API key and the AI feature your SaaS provider switched on last quarter are all in scope. Registry coverage, meaning systems in the register divided by systems found in an independent sweep, is the one metric that tests whether everything else measures the whole estate.
The engineering leader who owns the AI surface chairs it, with a security lead, whoever answers customer security questionnaires, and one product manager who owns an AI-facing feature.
Work through eight domains. Score by consensus rather than average, and where the room disagrees take the lower score and write down why, because that disagreement is usually the real finding.
Transfer the eight subtotals, total out of 160, then check for any domain below 10 out of 20. That domain is your finding regardless of what the headline number says.
Use the band-specific plan to sequence days 1-30, 31-60 and 61-90, then re-score quarterly and keep every scorecard so the trend is provable.
Two things. The evidence scale, because most models let you score intent, which is why most companies self-report as Defined and then fail an evidence-based review. And the sizing: this assumes no chief risk officer, no model validation function and no internal audit, because at 200 to 1,000 people you have none of those and a model that assumes them produces a plan you cannot execute.
No, and waiting for one means never starting. One of the eight domains is the inventory domain, so an incomplete register simply scores low and becomes your first finding. What matters is honesty about the denominator: score against your best estimate of the whole estate, not against the systems you happen to know about.
Four people. The engineering leader who owns the AI surface chairs it, plus a security or IT lead, whoever answers customer security questionnaires, and one product manager who owns an AI-facing feature. Score by consensus, and take the lower number wherever the room disagrees.
It tells you whether certification is worth starting. Certifying below roughly 112 is expensive theatre, because the audit surfaces the same gaps you would have found here for free. Above that, the assessment plus its framework crosswalk is a credible gap-assessment starting point and the evidence references become the index a certification body asks for.
Quarterly for the first year, then semi-annually once no domain sits below 14. Also after any material change: a new model provider, a first agentic deployment, an acquisition, or entry into a regulated vertical. Keep every completed scorecard, because the trend line is worth more in a customer conversation than any single score.
Yes. The AI Security Maturity Assessment uses the same scale and bands so the two results are comparable. Governance asks whether you are building the right things with the right oversight; security asks what happens when someone attacks what you already shipped. Most companies need both numbers and they are rarely the same.
Drop your details and we'll send AI Governance Maturity Assessment straight to your inbox - no spam, unsubscribe anytime.
Talk through your result with our engineering leads once you have scored it, and we will work through the weakest domain with you. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers