Count the estate first, since every row here is answered as a proportion rather than a yes. The denominator is every AI system currently taking production traffic, including the ones bought inside somebody else's software, and each of the sixteen rows asks what share of them clears the gate. One system scores out of a hundred; the estate score is the mean of those totals. What you finish holding is a spread: this many production-grade, this many coasting, this many that should not be taking traffic at all.
Take the mean of the per-system totals, then apply the floor rule before you say the band out loud. Each band is a statement about the estate as a portfolio, and each one changes what should be allowed to ship next.
Eighty and above, where the standard is the default rather than the exception and a new system inherits it without a campaign. Put the remaining effort into the AI you bought rather than built, since those entries score worst and appear in no engineering inventory. Rescore whenever the estate grows by a quarter, because growth is what dilutes a good mean.
Fifty to seventy-nine. The middle of the portfolio is sound and a minority is being subsidised by it, which is comfortable until somebody asks which of them a customer actually touches. Rank by reach rather than by score, and close release and reversibility and then operation and cost on the customer-facing systems first.
Under fifty, or any estate with two systems below the floor. Most of what is live cannot be changed safely, so stop adding systems, pick the four with the widest reach and take those through all four domains before anything else ships. The logistics group sat here at a mean of 68, which is the clearest argument for reading the floor first.
The count is the denominator for all sixteen rows, so build it from sign-on records, egress logs, expense lines and provider invoices, and write down what reconciliation added and when.
B3 carries 7 points and applies only to retrieval-backed systems. Count those separately, report recall at k and answerability apart from answer quality, and put the smaller denominator on the sheet.
D1 is worth 8 points, and a trace without document identifiers cannot tell you whether retrieval or generation failed. Most estates score this row on logging they already pay for.
A4 is worth 4 points and asks for a decision point with the authority to switch something off. Zero retirements across a large estate is a finding in its own right.
The estate, meaning every AI system currently taking production traffic, including AI features inside software you bought. No row is answered yes or no. Each one asks what share of the live systems clears it, so the sheet returns a distribution instead of a verdict on one build.
Definition and ownership takes 20, evaluation evidence 30, release and reversibility 25, and operation and cost 25, across sixteen rows weighted between 4 and 10 points each. The domains come from the four gates in the P12 production standard, and the weights follow what decides whether a live system can be changed safely.
It depends entirely on the shape underneath it. A logistics group reached exactly that across twelve systems, with four at or above 80 and two at 24 and 19, both customer-facing and neither holding an evaluation suite, a trace or a rehearsed rollback. Apply the floor rule and the band drops.
Both, and that is the check. The row score uses the proportion clearing the gate, the estate score is the mean of the per-system totals, and the two routes have to agree. Disagreement almost always means the denominator is wrong, which is the most useful thing the exercise finds in its first hour.
Roughly two weeks to measure the estate rather than the flagship, and most of that is reconciling the count against sign-on records, egress logs and provider invoices. The scoring itself is a session. Record the date, the denominator and who was in the room, since the next score is read against this one.
Treat it as an engineering problem rather than a scoring one. A row failing on two systems is a remediation list; the same row failing on nine is a missing platform capability, and Production AI: An Engineering Reference holds those controls by family. Fix it once in the pipeline and the proportion moves everywhere.
This one measures an estate. Benchmarking Retrieval Architecture scores a single retrieval architecture, on the argument that answer quality hides whether retrieval or generation failed, which is a failure this sheet only counts. Score the estate first, then take the low systems apart one at a time.
Drop your details and we'll send Benchmarking Production AI straight to your inbox - no spam, unsubscribe anytime.
Bring the inventory, including the entries nobody in engineering owns, and sit down with our engineering leads. We count the estate first, work the sixteen rows as proportions, and hand back the distribution with the systems under the floor named. SECTION 7 - FAQ - 5 to 8 questions
Book an estate scoring session