Classical application security rests on a boundary between instructions and data. Code is trusted, input is not, and the job is keeping them apart. A language model dissolves that boundary by design, because the instruction channel is the data channel. Everything difficult about AI security follows from that one property, and every control plane here exists to rebuild a boundary the architecture no longer provides for free. This is not a survey of what could go wrong. It is the five attack paths that actually reach a mid-market company with LLM features in production, and the six planes that answer them.
This architecture rests on three positions. If you disagree with any of them the design changes, so they are worth arguing about before you build.
Provider safety training reduces casual misuse. It is not a boundary, it changes without notice, and it is not contractually guaranteed to you. Designing as though it were means your security posture is set by somebody else's release calendar, and you will discover the change when an evaluation fails rather than when a notice arrives.
Design as though an attacker has it, because extraction is cheap and your prompt is one clever turn away from disclosure. Any control whose strength depends on the prompt staying hidden is already defeated. No secret, credential or authorisation decision should live there.
not the architecture
They raise attacker cost against known patterns and give you useful telemetry. They do not constrain what a credential can do, which is where severity actually lives. A guardrail purchased before tool scopes are bounded creates coverage over an unbounded action surface, which is worse than no coverage because it ends the conversation.
Build the tool registry, narrow every agent scope, issue per-service credentials, and set rate, token and spend ceilings. Two to three engineer-weeks, and it removes the only irreversible path.
Schema validation and contextual encoding at every output boundary, egress allowlisting, retrieval authorisation at query time and tenant-partitioned vector indexes.
Rearchitect prompt assembly for instruction and data separation, tag provenance on retrieved content, add approval gates on irreversible actions, then test injection through the real retrieval path.
Model-layer telemetry and detections, AI scenarios in the incident runbook, shadow AI discovery wired to SSO and spend data, and confirmation that a simulated attack pages a human.
No, and buying one before tool scopes are bounded is money spent in the wrong order. Guardrails are a worthwhile detection layer that raise attacker cost against known patterns. They do not constrain what a credential can do. Containment first because it works whether or not you detected the attack, detection second because knowing is worth paying for.
Architecturally, and imperfectly. Separate retrieved content from instructions in prompt assembly, tag provenance, sanitise untrusted content before it enters context, isolate context per tenant and session, and test through the real retrieval and tool paths. These bound what a successful injection achieves. Filtering on top reduces how often one succeeds.
Two, and they compound. Cross-tenant leakage if authorisation is enforced above the retrieval layer rather than at query time, which is reportable rather than merely embarrassing. And indirect injection, because a customer-uploaded document is attacker-controllable text your system reads as instruction. Partition indexes by tenant and filter by user within the partition.
Usually not, and generalised reassurance is the most common expensive reason teams do it. Self-hosting trades a vendor's security team for your own operational burden, and the result is typically less well patched. Choose it for a specific named constraint such as hard data residency, not as a posture.
Roughly 68 to 212 engineer-days depending on how much is retrofit rather than designed in, across four phases over three to six months. The first phase is two to three engineer-weeks and removes the irreversible path, which is why it should be asked for separately.
One plane is written for them and carries the heaviest weighting. Enumerate every callable tool with its scope and justification, read-only by default, separate approval for write and transact scopes, human approval before irreversible actions with real context, bounded blast radius, and a tested emergency stop. Treat every new agent as a first-phase event.
Drop your details and we'll send AI Security Reference Blueprint straight to your inbox - no spam, unsubscribe anytime.
Instruction separation and retrieval authorisation are properties of how you assemble context. Talk through your architecture with our engineering leads. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers