AI security costing goes wrong differently from governance costing. Governance cases fail for lack of a counterfactual. Security cases fail because they price the project and ignore the tax, meaning the latency, tokens, storage and compute that every defensive control adds to every request, forever. A guardrail that costs nothing in licence and 180 milliseconds plus 400 tokens per call is not free. At ten million calls a year it is a line item larger than the engineering that built it, and it arrives without a purchase order.
Three findings from the model change how most teams sequence their spend, and all three run against the intuitive order.
Instruction and data separation, retrieval authorisation, schema validation and scoped tool permissions together cost under $12,000 a year at nine million calls, and they address the two highest-severity attack paths. Guardrail products cost three to nine times as much and address the path that is hardest to close by any means.
Bounding agent scopes costs about $9,750 and removes irreversible production change. Rebuilding output and retrieval boundaries costs about $24,000 and removes cross-tenant breach. Hardening context costs about $21,000. The severity ratios differ by more than an order of magnitude while the costs differ by less, which is why ordering matters more than the total.
These ranges assume adding controls to a system already in production. If retrieval authorisation was never designed in, the boundary phase carries a re-index and a data migration adding fifteen to thirty engineer-days on its own. The same control costs roughly a third as much designed in rather than retrofitted.
Get calls per year, average tokens per call and blended cost per million tokens from the provider console, then project three years out.
Apply the per-control latency and token overheads to your own volume. This is the pool that decides vendor versus self-hosted, and it is usually missing entirely.
Eight priorities with typical effort and the reason for each position. Agent tool scoping first, guardrail purchases seventh, which is the inverse of most budgets.
Under $10,000 to remove the only irreversible attack path. Presented alone it is a decision; buried in a $233,500 programme it waits for the next planning cycle.
Roughly 68% in year one, converging toward 37% by year three as volume grows. That ratio is the most useful single planning figure in the model, because it scales with your business rather than your headcount and it argues for building controls before volume arrives rather than after.
Buy below roughly fifteen million calls a year, host above roughly forty million, and decide deliberately in between. The more important point is ordering: a guardrail bought before agent tool scopes are bounded creates the appearance of coverage over an unbounded action surface.
Typically $25,000 to $85,000 depending on scope and surface count. Book it at the end of the boundary phase rather than after everything ships, because findings that arrive while the architecture is still being designed change it, while the same findings six months later become a backlog.
The first phase, at roughly $9,750 and two to three engineer-weeks. Enumerate every agent tool and credential, revoke write access not justified in writing, issue per-service credentials, and set rate and spend ceilings. It is the cheapest phase and removes the only failure that destroys rather than leaks.
Yes, materially. The first phase shrinks, the largest exposure line disappears, and the agent override rules stop applying. Your remaining concentration is output handling and cross-tenant retrieval, both conventional problems reached through a novel path and both comparatively cheap to close.
Use the ratio and the counterfactual together. Security at 68% of inference spend sounds high until it sits beside $207,000 to $434,000 of annualised status-quo exposure, most of it deal friction rather than hypothetical breach cost. Then split the first phase out as a small, decisive ask.
Drop your details and we'll send AI Security Cost Calculator straight to your inbox - no spam, unsubscribe anytime.
Work through your own volume and architecture with our engineering leads, and we will tell you where the crossover sits. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers