Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

Benchmarking Agentic Systems.

Most maturity assessments hand back a number for the organisation, and that number is exactly what conceals the reconciliation agent holding a standing write token on the ledger with no record of what it did last Tuesday. This instrument scores one deployed agent, named and credentialed, and nothing wider than that. Seventeen rows across four domains, one hundred points, and each row pays its whole weight or none of it depending on whether the artefact or the measurement is in front of you. Then you start a fresh sheet for the next agent.

In depth

A Fleet Average Describes No Agent And Hides The One That Will Break.

01

Score the agent, not the organisation that owns it.

A maturity number for the company tells you agentic work is being taken seriously somewhere; it will not tell you that one agent can write to the ledger and cannot show what it did last Tuesday. Run the sheet once per agent, date it, record who was in the room, and compare the next score for that agent against this one.

In shortcompare the next score for that agent against this one
02

Points are distributed by what stops an agent, not by build effort.

Boundary and authorisation carries 30, evaluation and reliability 25, observability and recoverability 25, and economics under load 20. The orchestration everybody is proud of earns less here than the identity record nobody bothered to write down. Domains A and C get settled while the first tool is being wired up, and every cheap fix in them stops being cheap once the agent is live.

In shortDomains A and C get settled while the first tool is…
03

Forty-eight is what a working agent typically earns.

An order exceptions agent at a wholesale distributor, seven months live, four tools with write scope on two of them and roughly 900 invocations a day, reached 48 out of 100 and landed in the bottom band. It failed on the plan trace and on reversal, so when one run issued 340 duplicate credit notes in eleven minutes nobody could say which run had done it and recovery meant a database restore. Three sibling agents in the same estate scored 72, 68 and 74, and the four-agent mean of 65.5 describes none of them.

In short5 describes none of them
04

A control that holds only in staging earns nothing.

The row pays when the artefact is on the table or the measurement can be pointed at while the meeting is still running, and a stated intention pays zero. There is no partial credit anywhere on the sheet. An agent that has never been tested under load then scores what it has earned rather than what it was promised.

In shortAn agent that has never been tested under load then…
The detail

Three Bands, And What Each One Permits This Agent To Do Next.

Total the four domains and read the band against the agent you just scored, not against the estate it sits in. The band is a permission statement rather than a grade, and it answers one question: how much rope this agent keeps until the failed rows close.

Zone · 01

Runs unattended

Seventy-eight and above. Widen its scope deliberately, one tool or one action class at a time, and rescore on every widening rather than once a year. Sampling its output weekly is enough at this point. The trap here is scope creep arriving through a tool registry change that nobody treated as a release, so the score ages quietly.

Zone · 02

A hand on the stop

Fifty-two to seventy-seven, which is where most live agents sit. Hold the current scope exactly where it is, close the failed rows in Domain A and Domain C first, and do not hand it another write tool until both are shut. Adding a tool to an agent in this band is how a contained problem becomes an incident with a customer on the other end of it.

Zone · 03

Should not hold these credentials

Under fifty-two. Cut the agent back to read-only, or to proposing actions a human executes, and do it before the next sprint rather than after the next board meeting. Fix the boundary and the plan trace, then restore write scope one tool at a time. The worked example landed here at 48 while its team, its users and its dashboard all believed it was working.

By the numbers

The figures that make it a board-level conversation.

One agent
the unit scored here, never a department, a platform or a programme
40%
of agentic AI projects will be cancelled by the end of 2027
92%
of AI-related breaches hit organisations with no AI access controls in place
Inside the report

What you'll take away.

01

Step 1 - Write down the agent's record before you score a single row

Identifier, named owner, autonomy boundary, every tool it can call and which of those write or spend, plus credential scope and expiry. Half the failed rows become obvious while you fill this in.

02

Step 2 - Fire the stop against a live run and date the record

C2 is worth 6 points and needs one production run halted between two tool calls. Write down the elapsed seconds and the state the system was left in, since an untested stop is a button.

03

Step 3 - Build the reversal table before you need it on a Sunday

C3 carries 6 points. Every write tool gets a named undo path, a compensating transaction or a written acceptance that it cannot be undone, with one reversal executed and timed on a working day.

04

Step 4 - Divide last month's spend by tasks actually completed

D1 is worth 6 points and most teams report cost per call instead. Retries, abandoned runs and human rework belong inside the number, and the tail matters more than the mean, so state the 99th percentile next to it.

Questions

Frequently asked.

What exactly does this scorecard measure?

One deployed agent, named in the orchestration configuration and holding its own credentials. Not the organisation, not the platform it runs on, and not the team that built it. If you have five agents in production you fill in five sheets, since the average across them is the thing that hides the dangerous one.

Where do the hundred points sit?

Boundary and authorisation takes 30, evaluation and reliability 25, observability and recoverability 25, and economics under load 20, across seventeen rows weighted between 4 and 9 points each. The split follows what determines whether an agent can keep acting, rather than what was hardest or most expensive to build.

Is 48 out of 100 a poor result?

It is a normal one for an agent that works. The wholesale distributor's order exceptions agent had been live seven months, ran around 900 invocations a day and nobody had complained about it. What it lacked was a plan trace and a reversal path, which is why one bad run cost a database restore.

Can a row take part of its weight?

No. Full weight or zero, decided on whether the artefact is on the table or the measurement can be pointed at during the meeting. Partial credit turns seventeen sharp questions into a mood reading, and a mood reading returns a comfortable number nobody schedules work against.

We run twelve agents. Does that mean twelve sheets?

Yes, and start with the one holding the widest credential scope rather than the newest. Scoring is roughly two hours per agent once the record block is filled in. Keep the sheets side by side and read the spread, since the lowest agent is the one that sets your actual exposure.

How does this sit alongside your other agentic resources?

This one measures an agent you already run. Agentic Systems: What Buyers Should Ask puts the same demands to a vendor while you are still choosing one, and Agentic Systems: An Engineering Reference holds the control families and the artefact each one produces, which is where a failed row gets fixed.

Who should fill it in, and how long does it take?

Head of AI, CISO or platform lead, with whoever configured the agent's tools and whoever holds its identity record. Two hours with the evidence column open. Record the date and the names, then give every failed row an owner and a date before anyone leaves the room.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Benchmarking Agentic Systems straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Pick the agent with the widest credential scope, not the one you like most.

Bring whoever configured its tools and sit down with our engineering leads. We work the seventeen rows together, write down every artefact that turns out not to exist, and rank the fixes by exposure removed per pound. SECTION 7 - FAQ - 5 to 8 questions

Book an agent scoring session