Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

Agentic Systems: What Buyers Should Ask.

A chatbot that gets it wrong hands you a sentence you can disregard. An agent that gets it wrong issues a refund, deletes a record, mails a client, and then takes four more steps that assumed the first one worked. So the due diligence changes shape. You are no longer assessing answer quality; you are assessing an action space somebody else designed, with your own credentials running inside it. These eight themes put that design on the table, and each one is written to be asked in writing and answered with an artefact.

In depth

Three Properties Set The Blast Radius, And All Three Are Architectural.

01

Ask where the autonomy boundary is enforced, not where it is described.

The reply you will hear is that the agent is instructed never to take irreversible actions without checking, and that the model follows that instruction reliably. That is a hint with a confidence interval attached to it. What holds instead is a policy layer that refuses the call at the point of execution whatever the plan says, plus goal changes and sub-agent delegation as enumerated capabilities you can switch off per deployment. A spawned agent should inherit a narrower scope than its parent and never a wider one.

In shortA spawned agent should inherit a narrower scope than…
02

Your own audit log settles the identity question.

Ask which identity appears in your CRM record when the agent writes to it, and a weak vendor names the integration account set up during onboarding. The defensible version is a distinct machine identity per agent and per environment, with the triggering human recorded beside it, and short-lived tokens minted per run and scoped to the single operation. Then test the stale-authority case: a user starts a run and loses access to a system halfway through it. The next tool call should fail closed and the run should halt rather than complete under authority that outlived its grant.

In shortThe next tool call should fail closed and the run sh…
03

Memory is the third architectural property, and the least examined.

Vendors describe it as the agent learning your processes over time, which is a value proposition rather than a design. Ask instead what is written, by which step, at what scope - run, user or tenant - with a retention period per class and a way to inspect and delete an entry. The question underneath is whether content the agent read on Monday can become an instruction it obeys on Thursday. Memory read back into a plan is untrusted input, and it should carry provenance and a scope that stops a poisoned entry crossing a user or a tenant.

In shortMemory read back into a plan is untrusted input, and…
04

A success rate is a marketing number.

The engineering number is a failure taxonomy: wrong tool, right tool with wrong arguments, premature completion, silent partial completion, loop, unnecessary escalation, each with an observed frequency that moves between releases. Ask too how a run that finished cleanly and did the wrong thing gets counted, and whether it lands in a separate silent-failure class with a named detection method. On observability, hold out for the plan and not the activity timeline: reasoning steps, tool calls with arguments and returns, retries, cost lines, escalation points not taken. The runs that report success while acting wrongly are the ones that reach a customer.

In shortThe runs that report success while acting wrongly ar…
The detail

Reversal, The Bill And The Exit: Settle These While You Can Still Walk.

An agent is the first software you have bought that can raise its own bill without anyone deciding to let it, and the first whose accumulated configuration is written largely by the vendor. Reversal belongs in the same conversation, because it is designed in or it is not there.

Zone · 01

Compensating actions

Step four of five fails and three side effects are already out in the world. A weak vendor marks the run failed and raises an alert for your team to review. A strong one has a named compensating action per side-effecting step, or an explicit design-time flag saying the step cannot be undone, plus a published list of irreversible operations that require approval by default.

Zone · 02

A ceiling that stops

Usage pricing sounds fair until a run loops. Ask what you are billed for on retries, failed runs and internal reflection steps, named individually, and ask to see a month where something went wrong with the line pointed at. Then ask for a hard ceiling that halts execution rather than alerting afterwards, configurable per environment, and what the agent does when it hits one.

Zone · 03

What leaves with you

Termination is where the lock-in becomes visible. Make them enumerate what leaves: tool manifests and permission scopes, escalation thresholds, prompts and policies, memory contents, evaluation cases and results, and the full trace archive, plus an honest statement of what stays behind. Then ask what happens to runs still in flight on the day you switch it off, and look for a drain mode with a report of what was mid-execution.

By the numbers

The figures that make it a board-level conversation.

8 themes
covering autonomy, identity, memory, evidence, traces, rollback, pricing and exit
40%
of organisations control access to their AI models and data at all
ASI01-10
the OWASP Top 10 for Agentic Applications, published December 2025 and worth naming in the room
Inside the report

What you'll take away.

01

Step 1 - Run the trial in your estate, under credentials you issue

A vendor sandbox tests the demo. Your systems test the integration, the permission model and the four connectors nobody mentioned, and anything they will only show inside their own tenancy is a finding.

02

Step 2 - Kill a downstream system at step three of five

Then read what the agent did with the two completed steps. Compensating actions, idempotency and the escalation path all become visible in one deliberate failure, and it takes an afternoon to stage.

03

Step 3 - Plant a poisoned document in the corpus it reads

Whether an instruction hidden in content survives into memory and changes a later run is the cheapest test of ASI01 and ASI06 you will ever run. Check the next day's behaviour, not the same session.

04

Step 4 - Export everything on the last day of the trial

Traces, memory contents and the tool manifest, requested while you are still a prospect. The exit clause is worth exactly what that export produced, and the answer arrives within the hour or it does not.

Questions

Frequently asked.

Which theme should we put first?

Autonomy, identity and memory, in that order. All three are architectural, so they are settled before you arrive and cannot be fixed by a clause. A vendor who describes any of them in terms of prompt instructions has told you the boundary is advisory, and the rest of the conversation is decoration.

The vendor quotes a task success rate above 90%. What now?

Ask for the other document. A success rate averages away the run that completed cleanly and did the wrong thing, which is the run that reaches a customer. What you want is a named list of failure classes with observed frequencies per release: wrong tool, wrong arguments, premature completion, silent partial completion, loop, unnecessary escalation.

Is a vendor sandbox trial good enough?

No, and refusing to leave it is itself the finding. A sandbox proves the demo works on data they chose. Your environment proves the integration, the permission model and the connectors nobody mentioned in the deck. Supply your own cases too, including twenty that went wrong last quarter, and keep the results whatever they show.

What does a defensible answer on identity actually sound like?

A distinct machine identity per agent and per environment, short-lived tokens minted per run and scoped to one operation, and the triggering human recorded alongside. They should be able to say what the agent cannot do with the token it holds right now. One shared integration account behind every tool is the answer that fails.

How does this sit alongside your other agentic resources?

This one is for the choosing. Benchmarking Agentic Systems is the instrument for an agent you already run, scored one agent at a time against what it can actually show. Ask the eight themes before signature, then score the thing you bought once it has been live a quarter.

How long does this take to run properly?

Two meetings and a two-week proof of concept. The first meeting puts the eight themes on the table and collects what can be shown on screen. The second works through what was missing. The proof of concept runs on your systems, with your cases and one deliberate failure staged at step three.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Agentic Systems: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Send us the eight themes and hold us to the artefacts.

Bring the eight themes to a working session with our engineering leads. We answer each one with a file rather than a sentence, and tell you which of them a vendor can honestly refuse. SECTION 7 - FAQ - 5 to 8 questions

Book an agent platform review