Most maturity assessments hand back a number for the organisation, and that number is exactly what conceals the reconciliation agent holding a standing write token on the ledger with no record of what it did last Tuesday. This instrument scores one deployed agent, named and credentialed, and nothing wider than that. Seventeen rows across four domains, one hundred points, and each row pays its whole weight or none of it depending on whether the artefact or the measurement is in front of you. Then you start a fresh sheet for the next agent.
Total the four domains and read the band against the agent you just scored, not against the estate it sits in. The band is a permission statement rather than a grade, and it answers one question: how much rope this agent keeps until the failed rows close.
Seventy-eight and above. Widen its scope deliberately, one tool or one action class at a time, and rescore on every widening rather than once a year. Sampling its output weekly is enough at this point. The trap here is scope creep arriving through a tool registry change that nobody treated as a release, so the score ages quietly.
Fifty-two to seventy-seven, which is where most live agents sit. Hold the current scope exactly where it is, close the failed rows in Domain A and Domain C first, and do not hand it another write tool until both are shut. Adding a tool to an agent in this band is how a contained problem becomes an incident with a customer on the other end of it.
Under fifty-two. Cut the agent back to read-only, or to proposing actions a human executes, and do it before the next sprint rather than after the next board meeting. Fix the boundary and the plan trace, then restore write scope one tool at a time. The worked example landed here at 48 while its team, its users and its dashboard all believed it was working.
Identifier, named owner, autonomy boundary, every tool it can call and which of those write or spend, plus credential scope and expiry. Half the failed rows become obvious while you fill this in.
C2 is worth 6 points and needs one production run halted between two tool calls. Write down the elapsed seconds and the state the system was left in, since an untested stop is a button.
C3 carries 6 points. Every write tool gets a named undo path, a compensating transaction or a written acceptance that it cannot be undone, with one reversal executed and timed on a working day.
D1 is worth 6 points and most teams report cost per call instead. Retries, abandoned runs and human rework belong inside the number, and the tail matters more than the mean, so state the 99th percentile next to it.
One deployed agent, named in the orchestration configuration and holding its own credentials. Not the organisation, not the platform it runs on, and not the team that built it. If you have five agents in production you fill in five sheets, since the average across them is the thing that hides the dangerous one.
Boundary and authorisation takes 30, evaluation and reliability 25, observability and recoverability 25, and economics under load 20, across seventeen rows weighted between 4 and 9 points each. The split follows what determines whether an agent can keep acting, rather than what was hardest or most expensive to build.
It is a normal one for an agent that works. The wholesale distributor's order exceptions agent had been live seven months, ran around 900 invocations a day and nobody had complained about it. What it lacked was a plan trace and a reversal path, which is why one bad run cost a database restore.
No. Full weight or zero, decided on whether the artefact is on the table or the measurement can be pointed at during the meeting. Partial credit turns seventeen sharp questions into a mood reading, and a mood reading returns a comfortable number nobody schedules work against.
Yes, and start with the one holding the widest credential scope rather than the newest. Scoring is roughly two hours per agent once the record block is filled in. Keep the sheets side by side and read the spread, since the lowest agent is the one that sets your actual exposure.
This one measures an agent you already run. Agentic Systems: What Buyers Should Ask puts the same demands to a vendor while you are still choosing one, and Agentic Systems: An Engineering Reference holds the control families and the artefact each one produces, which is where a failed row gets fixed.
Head of AI, CISO or platform lead, with whoever configured the agent's tools and whoever holds its identity record. Two hours with the evidence column open. Record the date and the names, then give every failed row an owner and a date before anyone leaves the room.
Drop your details and we'll send Benchmarking Agentic Systems straight to your inbox - no spam, unsubscribe anytime.
Bring whoever configured its tools and sit down with our engineering leads. We work the seventeen rows together, write down every artefact that turns out not to exist, and rank the fixes by exposure removed per pound. SECTION 7 - FAQ - 5 to 8 questions
Book an agent scoring session