Agent pilots demo well and die in production, and the cause is almost never the model. This report works through the five failures that account for nearly every cancelled programme, the arithmetic of error compounding across a nine step chain, and the four gates a survivor passes.
The trap most platform teams walked into: hand the model broad credentials and a large tool catalogue, ship a nine step chain because the demo path was three, retrofit tests onto behaviour that is already contested, and report cost per token while the real figure is cost per resolved case.
What the surviving programmes do instead: spend the first two weeks building an evaluation harness that produces no demo, write the task boundary in one sentence, cut the chain to the fewest steps that finish the job, and give every tool a typed contract with its own credential.
If the job is triaging refunds under $200, the agent needs read access to the order record, one refund call with a hard ceiling and one escalation call. Six typed tools give you a state space you can describe to a security reviewer. A generic HTTP tool gives you none.
Real historical cases with graded outcomes, running in CI and reporting completion, tool call correctness, cost and latency per run. Without one, somebody changes a prompt, the tool descriptions and the retry policy in the same week, quality moves, and nobody can attribute it. That is adjusting prompts, not engineering.
Token spend counts attempts. The fully loaded cost of one completed, accepted task absorbs failed attempts, human review and the cleanup after a bad outcome. A programme reporting $0.04 per call can be spending $2.60 per resolved case once nine attempts and a reviewer are counted, and that gap is where budgets die.
Name the job, the decision the agent may take alone, and where it must stop. If it needs three clauses, the scope is too wide to test, and an untestable scope is the thing that gets cancelled in month nine.
A few hundred real historical cases with graded outcomes, including the awkward ones. About two weeks of work, no demo at the end of it, and the highest return decision in the programme, because every later change becomes attributable.
Track step count as a first class metric rather than an implementation detail. Going from ten steps to five at 95% per step moves completion from 60% to 77%, which is more reliability than a better model buys you.
Weekly completion, escalation rate and cost per accepted outcome against the manual baseline, modelled at ten times current volume. Checkpoint state after every step and key every side effecting call, so a retry cannot double charge and a failed run resumes at step seven.
At 95% per step, three steps complete 86% of the time and nine complete 63%, because per step reliability multiplies. The demo path is short and clean. Production adds steps, and each one takes a fixed percentage off the
end to end number.
Buy the substrate, build the judgement. Orchestration, tracing and evaluation plumbing are commodity. A vertical agent means the vendor owns your action space, escalation logic and failure taxonomy, which encode your operating rules while you carry the consequences.
Four questions, each answered with an artefact. Show the evaluation suite and its pass rate on cases you did not write. Show the action space tool by tool with permission scopes. Show cost per completed task at ten times our volume. Show a trace of step four of nine failing.
Yes, and it is worth more now than it was. Pull a few hundred real historical cases from your logs, grade the outcomes, and freeze them. That gives you the first attributable measurement of a change, which is exactly what defunded programmes could never produce.
The two measure different things. Token spend counts attempts, cost per completed task counts outcomes, so it absorbs retries, human review and post failure cleanup. A reported $0.04 per call can be $2.60 per resolved case,
and only the second figure behaves sensibly at ten times volume.
Drop your details and we'll send 40% Of Agent Projects Get Cancelled. Five Engineering Failures Explain Why straight to your inbox - no spam, unsubscribe anytime.
Bring the agent you cannot get past a security or finance review, and our engineering leads will show you which gate it fails.
Book a readiness review