Logiciel Contact Us
Success Stories Tech News Contact Us
whitepaper

40% Of Agent Projects Get Cancelled. Five Engineering Failures Explain Why.

Agent pilots demo well and die in production, and the cause is almost never the model. This report works through the five failures that account for nearly every cancelled programme, the arithmetic of error compounding across a nine step chain, and the four gates a survivor passes.

In depth

The Demo Completes Every Time. Production Completes 60%.

01

The trap most platform teams walked into: hand the model broad credentials and a large tool catalogue, ship a nine step chain because the demo path was three, retrofit tests onto behaviour that is already contested, and report cost per token while the real figure is cost per resolved case.

02

What the surviving programmes do instead: spend the first two weeks building an evaluation harness that produces no demo, write the task boundary in one sentence, cut the chain to the fewest steps that finish the job, and give every tool a typed contract with its own credential.

The detail

What Separates The Agents That Survive.

Zone · 01

A Bounded Action Space

If the job is triaging refunds under $200, the agent needs read access to the order record, one refund call with a hard ceiling and one escalation call. Six typed tools give you a state space you can describe to a security reviewer. A generic HTTP tool gives you none.

Zone · 02

A Frozen Evaluation Suite

Real historical cases with graded outcomes, running in CI and reporting completion, tool call correctness, cost and latency per run. Without one, somebody changes a prompt, the tool descriptions and the retry policy in the same week, quality moves, and nobody can attribute it. That is adjusting prompts, not engineering.

Zone · 03

Cost Per Completed Task

Token spend counts attempts. The fully loaded cost of one completed, accepted task absorbs failed attempts, human review and the cleanup after a bad outcome. A programme reporting $0.04 per call can be spending $2.60 per resolved case once nine attempts and a reviewer are counted, and that gap is where budgets die.

By the numbers

The figures that make it a board-level conversation.

40%
of agentic AI projects will be scrapped before the end of 2027, on escalating cost, unclear business value and inadequate risk controls
95%
of enterprise AI pilots produce no measurable profit impact. Agent pilots inherit that deployment gap with considerably more moving parts
130
vendors Gartner judges genuinely agentic, out of thousands
Inside the report

What you'll take away.

01

Write the task boundary in one sentence

Name the job, the decision the agent may take alone, and where it must stop. If it needs three clauses, the scope is too wide to test, and an untestable scope is the thing that gets cancelled in month nine.

02

Build the harness before the agent

A few hundred real historical cases with graded outcomes, including the awkward ones. About two weeks of work, no demo at the end of it, and the highest return decision in the programme, because every later change becomes attributable.

03

Cut the chain and count the steps

Track step count as a first class metric rather than an implementation detail. Going from ten steps to five at 95% per step moves completion from 60% to 77%, which is more reliability than a better model buys you.

04

Report on outcomes, not tokens

Weekly completion, escalation rate and cost per accepted outcome against the manual baseline, modelled at ten times current volume. Checkpoint state after every step and key every side effecting call, so a retry cannot double charge and a failed run resumes at step seven.

Questions

Frequently asked.

Our agent works in the demo, so why would it fail at nine steps?

At 95% per step, three steps complete 86% of the time and nine complete 63%, because per step reliability multiplies. The demo path is short and clean. Production adds steps, and each one takes a fixed percentage off the
end to end number.

Should we buy a vertical agent instead of building one ourselves?

Buy the substrate, build the judgement. Orchestration, tracing and evaluation plumbing are commodity. A vertical agent means the vendor owns your action space, escalation logic and failure taxonomy, which encode your operating rules while you carry the consequences.

What should we ask a vendor claiming their product is agentic?

Four questions, each answered with an artefact. Show the evaluation suite and its pass rate on cases you did not write. Show the action space tool by tool with permission scopes. Show cost per completed task at ten times our volume. Show a trace of step four of nine failing.

We are already live, so is a harness still worth two weeks now?

Yes, and it is worth more now than it was. Pull a few hundred real historical cases from your logs, grade the outcomes, and freeze them. That gives you the first attributable measurement of a change, which is exactly what defunded programmes could never produce.

Is cost per completed task not just token spend with extra steps?

The two measure different things. Token spend counts attempts, cost per completed task counts outcomes, so it absorbs retries, human review and post failure cleanup. A reported $0.04 per call can be $2.60 per resolved case,
and only the second figure behaves sensibly at ten times volume.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send 40% Of Agent Projects Get Cancelled. Five Engineering Failures Explain Why straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Prove one narrow agent.

Bring the agent you cannot get past a security or finance review, and our engineering leads will show you which gate it fails.

Book a readiness review