Nobody decided that the guardrail should fail open. Nobody decided the agent would hold a shared service credential, or that prompt logs would land in the observability store forty engineers can query. Those are defaults that arrived with a library, a tutorial or a deadline and were never examined. In governance an unexamined decision produces an inconsistency. In security it produces a control that looks present and is not, which is worse than an absent control because it ends the conversation. Every decision here is one an incident review typically surfaces after the fact.
Most of the ten decisions are cheap to change later. These three are not, and each has a default that is right for almost every mid-market case.
with degraded mode
Output validation, retrieval authorisation and tool permission checks fail closed always. Input guardrails fail closed on higher-risk surfaces. The option most teams do not have is degraded mode, where the model still responds but with tools and retrieval disabled and a restricted capability set, which removes the false choice between full function and no function during an incident.
not either
Post-filtering alone is not an implementation, it is an unresolved finding, because unauthorised content has already entered the process and leaks through ranking, counts and timing. Query-time metadata filtering is the baseline and per-tenant partitioning is the isolation. You need both, because partitioning stops a cross-customer leak and filtering stops a within-customer privilege problem.
Placement matters less than content. A prompt reading that the agent wants to send an email produces rubber-stamping within a week. Show the recipient, the full content, what triggered it and what happens next. Then measure: above roughly 95% approval the gate is decorative, and you should narrow what reaches it or admit what it is.
If the answer is that somebody would have to check, the decision was inherited rather than made. Start there, because it is the fastest diagnostic you have.
Fail behaviour, the agent credential model, the threshold for pulling a system offline, and retrieval authorisation, which is the only one that gets materially more expensive with time.
Options considered, what was chosen and why, who disagreed, and the condition that should reopen it. Two or three sentences, because longer entries do not get written.
Take the dependency down, trigger the emergency stop, exercise degraded mode. A security decision that has not been exercised is a statement of intent, not a control.
Closed on anything that enforces a boundary. Output validation, retrieval authorisation and tool permission checks fail closed always. Input guardrails fail closed on higher-risk surfaces. Where availability genuinely matters, use degraded mode with tools and retrieval disabled rather than choosing between full function and none. Then test it by taking the dependency down.
Per-agent identity always, plus user-delegated wherever a human initiated the request. The intersection is the narrowest permission set available and the only model that survives a customer asking whether your agent could have read data belonging to a user who did not request it. A shared service account is unacceptable for any write scope.
No on its own. Partition by tenant and filter by user within the partition. They fail differently: partitioning prevents a cross-customer leak, filtering prevents a within-customer privilege problem. Post-filtering alone should be treated as an unresolved finding, because unauthorised content has already entered the process.
Yes, with controls. Full capture in a separate store, restricted access with a recorded reason, bounded retention of 30 to 90 days on lower-risk surfaces, and the access log itself reviewed. Treat them like your production database. The failure to avoid is full capture landing in a general lake that forty engineers can query.
Yes, under six conditions: verified source with recorded provenance, safe serialisation formats only, checksum or signature verified, scanned before entering any environment touching production, first load in isolation with no credentials or egress, and licence reviewed for commercial use. Keep the serialisation rule absolute.
Four conditions, decided in advance and delegated to the on-call commander with notification after rather than approval before. Evidence of data exposure across a tenant or user boundary. An agent action with real-world effect whose path is not understood. A reproducibly working injection technique against production. Or model behaviour changing in a way that breaks a control.
Drop your details and we'll send AI Security Decision Framework straight to your inbox - no spam, unsubscribe anytime.
Decide it before you index at scale. Work through the four that cannot wait with our engineering leads. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers