A chatbot that gets it wrong hands you a sentence you can disregard. An agent that gets it wrong issues a refund, deletes a record, mails a client, and then takes four more steps that assumed the first one worked. So the due diligence changes shape. You are no longer assessing answer quality; you are assessing an action space somebody else designed, with your own credentials running inside it. These eight themes put that design on the table, and each one is written to be asked in writing and answered with an artefact.
An agent is the first software you have bought that can raise its own bill without anyone deciding to let it, and the first whose accumulated configuration is written largely by the vendor. Reversal belongs in the same conversation, because it is designed in or it is not there.
Step four of five fails and three side effects are already out in the world. A weak vendor marks the run failed and raises an alert for your team to review. A strong one has a named compensating action per side-effecting step, or an explicit design-time flag saying the step cannot be undone, plus a published list of irreversible operations that require approval by default.
Usage pricing sounds fair until a run loops. Ask what you are billed for on retries, failed runs and internal reflection steps, named individually, and ask to see a month where something went wrong with the line pointed at. Then ask for a hard ceiling that halts execution rather than alerting afterwards, configurable per environment, and what the agent does when it hits one.
Termination is where the lock-in becomes visible. Make them enumerate what leaves: tool manifests and permission scopes, escalation thresholds, prompts and policies, memory contents, evaluation cases and results, and the full trace archive, plus an honest statement of what stays behind. Then ask what happens to runs still in flight on the day you switch it off, and look for a drain mode with a report of what was mid-execution.
A vendor sandbox tests the demo. Your systems test the integration, the permission model and the four connectors nobody mentioned, and anything they will only show inside their own tenancy is a finding.
Then read what the agent did with the two completed steps. Compensating actions, idempotency and the escalation path all become visible in one deliberate failure, and it takes an afternoon to stage.
Whether an instruction hidden in content survives into memory and changes a later run is the cheapest test of ASI01 and ASI06 you will ever run. Check the next day's behaviour, not the same session.
Traces, memory contents and the tool manifest, requested while you are still a prospect. The exit clause is worth exactly what that export produced, and the answer arrives within the hour or it does not.
Autonomy, identity and memory, in that order. All three are architectural, so they are settled before you arrive and cannot be fixed by a clause. A vendor who describes any of them in terms of prompt instructions has told you the boundary is advisory, and the rest of the conversation is decoration.
Ask for the other document. A success rate averages away the run that completed cleanly and did the wrong thing, which is the run that reaches a customer. What you want is a named list of failure classes with observed frequencies per release: wrong tool, wrong arguments, premature completion, silent partial completion, loop, unnecessary escalation.
No, and refusing to leave it is itself the finding. A sandbox proves the demo works on data they chose. Your environment proves the integration, the permission model and the connectors nobody mentioned in the deck. Supply your own cases too, including twenty that went wrong last quarter, and keep the results whatever they show.
A distinct machine identity per agent and per environment, short-lived tokens minted per run and scoped to one operation, and the triggering human recorded alongside. They should be able to say what the agent cannot do with the token it holds right now. One shared integration account behind every tool is the answer that fails.
This one is for the choosing. Benchmarking Agentic Systems is the instrument for an agent you already run, scored one agent at a time against what it can actually show. Ask the eight themes before signature, then score the thing you bought once it has been live a quarter.
Two meetings and a two-week proof of concept. The first meeting puts the eight themes on the table and collects what can be shown on screen. The second works through what was missing. The proof of concept runs on your systems, with your cases and one deliberate failure staged at step three.
Drop your details and we'll send Agentic Systems: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.
Bring the eight themes to a working session with our engineering leads. We answer each one with a file rather than a sentence, and tell you which of them a vendor can honestly refuse. SECTION 7 - FAQ - 5 to 8 questions
Book an agent platform review