Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

Benchmarking Production AI.

Count the estate first, since every row here is answered as a proportion rather than a yes. The denominator is every AI system currently taking production traffic, including the ones bought inside somebody else's software, and each of the sixteen rows asks what share of them clears the gate. One system scores out of a hundred; the estate score is the mean of those totals. What you finish holding is a spread: this many production-grade, this many coasting, this many that should not be taking traffic at all.

In depth

A Mean Of Sixty-Eight Can Describe A Healthy Estate Or A Dangerous One.

01

Estates fail unevenly, and the average is where that unevenness goes to hide.

A logistics group with twelve live systems scored 68, which reads like a competent programme carrying a little debt. Under that mean sat four systems at 91, 88, 84 and 80, six between 67 and 78, and two at 24 and 19. The mean described none of the twelve, and the two at the bottom were the ones a customer could reach.

In shortThe mean described none of the twelve, and the two a…
02

The weights follow what makes a live system changeable.

Definition and ownership carries 20 points, evaluation evidence 30, release and reversibility 25, and operation and cost 25. Nothing is weighted by how hard it was to build, which is why the retrieval pipeline everybody argued about earns less here than the rehearsed rollback nobody scheduled. A row pays its weight only where every counted system clears it.

In shortA row pays its weight only where every counted syste…
03

Work every row twice, and make the two routes agree.

Count how many live systems clear the gate and score the row at its weight times that proportion, then keep a per-system tally and take the mean of the totals. The two numbers have to match, and where they do not, the inventory is usually wrong rather than the arithmetic. Most estates are larger than the inventory says, so an undercount flatters every score on the sheet.

In shortan undercount flatters every score on the sheet
04

Read the floor before you read the mean.

Any live system under 50 drops the estate one band, and two or more put it in the bottom band whatever the average says. That rule turned the logistics group's 68 from carrying a tail into carrying passengers, on the strength of a claims summariser at 24 and a pricing recommender at 19. Count the systems under 50 separately, since that is the number the mean will not show you.

In shortsince that is the number the mean will not show you
The detail

Where The Spread Puts You, And What It Says About The Next Build.

Take the mean of the per-system totals, then apply the floor rule before you say the band out loud. Each band is a statement about the estate as a portfolio, and each one changes what should be allowed to ship next.

Zone · 01

Carries itself

Eighty and above, where the standard is the default rather than the exception and a new system inherits it without a campaign. Put the remaining effort into the AI you bought rather than built, since those entries score worst and appear in no engineering inventory. Rescore whenever the estate grows by a quarter, because growth is what dilutes a good mean.

Zone · 02

Carrying a tail

Fifty to seventy-nine. The middle of the portfolio is sound and a minority is being subsidised by it, which is comfortable until somebody asks which of them a customer actually touches. Rank by reach rather than by score, and close release and reversibility and then operation and cost on the customer-facing systems first.

Zone · 03

Carrying passengers

Under fifty, or any estate with two systems below the floor. Most of what is live cannot be changed safely, so stop adding systems, pick the four with the widest reach and take those through all four domains before anything else ships. The logistics group sat here at a mean of 68, which is the clearest argument for reading the floor first.

By the numbers

The figures that make it a board-level conversation.

A distribution
what the sheet returns: this many production-grade, this many coasting, this many that should not take traffic
95%
of enterprise generative AI pilots show no measurable effect on profit
7.2%
fall in delivery stability for each 25% rise in AI adoption
Inside the report

What you'll take away.

01

Step 1 - Reconcile the inventory before you score anything

The count is the denominator for all sixteen rows, so build it from sign-on records, egress logs, expense lines and provider invoices, and write down what reconciliation added and when.

02

Step 2 - State the sub-denominator on the retrieval row

B3 carries 7 points and applies only to retrieval-backed systems. Count those separately, report recall at k and answerability apart from answer quality, and put the smaller denominator on the sheet.

03

Step 3 - Count the traces that actually carry retrieved context

D1 is worth 8 points, and a trace without document identifiers cannot tell you whether retrieval or generation failed. Most estates score this row on logging they already pay for.

04

Step 4 - Record how many systems the quarterly review has retired

A4 is worth 4 points and asks for a decision point with the authority to switch something off. Zero retirements across a large estate is a finding in its own right.

Questions

Frequently asked.

What is the unit being scored here?

The estate, meaning every AI system currently taking production traffic, including AI features inside software you bought. No row is answered yes or no. Each one asks what share of the live systems clears it, so the sheet returns a distribution instead of a verdict on one build.

How are the hundred points split?

Definition and ownership takes 20, evaluation evidence 30, release and reversibility 25, and operation and cost 25, across sixteen rows weighted between 4 and 10 points each. The domains come from the four gates in the P12 production standard, and the weights follow what decides whether a live system can be changed safely.

Our mean is 68. Is that a good estate?

It depends entirely on the shape underneath it. A logistics group reached exactly that across twelve systems, with four at or above 80 and two at 24 and 19, both customer-facing and neither holding an evaluation suite, a trace or a rehearsed rollback. Apply the floor rule and the band drops.

Do we score proportions or individual systems?

Both, and that is the check. The row score uses the proportion clearing the gate, the estate score is the mean of the per-system totals, and the two routes have to agree. Disagreement almost always means the denominator is wrong, which is the most useful thing the exercise finds in its first hour.

How long does scoring an estate take?

Roughly two weeks to measure the estate rather than the flagship, and most of that is reconciling the count against sign-on records, egress logs and provider invoices. The scoring itself is a session. Record the date, the denominator and who was in the room, since the next score is read against this one.

What do we do about a row that fails across most of the estate?

Treat it as an engineering problem rather than a scoring one. A row failing on two systems is a remediation list; the same row failing on nine is a missing platform capability, and Production AI: An Engineering Reference holds those controls by family. Fix it once in the pipeline and the proportion moves everywhere.

Where does this sit next to your other benchmarks?

This one measures an estate. Benchmarking Retrieval Architecture scores a single retrieval architecture, on the argument that answer quality hides whether retrieval or generation failed, which is a failure this sheet only counts. Score the estate first, then take the low systems apart one at a time.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Benchmarking Production AI straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Bring the whole list, including the ones built by people outside your team.

Bring the inventory, including the entries nobody in engineering owns, and sit down with our engineering leads. We count the estate first, work the sixteen rows as proportions, and hand back the distribution with the systems under the floor named. SECTION 7 - FAQ - 5 to 8 questions

Book an estate scoring session