A conventional breach has a perimeter you can reason about. An injection that succeeded in a retrieval pipeline may have touched one conversation or every conversation that retrieved the same document, and answering that is the investigation rather than the trigger. That asymmetry drives the central rule of this runbook: contain first, scope second. Every playbook is ordered accordingly, and every one carries a do-not column populated mostly with the instinct to establish scope before acting.
Three differences change how the first hour runs. Each one is a place where the AI version of an incident diverges from the conventional one.
not the complaint
One user noticed, and the retrieval logs will usually show the same injected content reached many more sessions. The difference between those two numbers is the difference between an internal fix and a notification obligation. Every playbook scopes from logs, and the do-not column exists largely to interrupt the instinct to establish scope before containing.
chosen in advance
Full disable for the highest severity or where you cannot characterise what the system is doing. Scope reduction where the affected path is narrow and identified. And degraded mode, where the model responds with tools and retrieval disabled, which is the option most teams lack and most need because it removes the argument that disabling the product is intolerable.
Seven post-incident questions, each with what a good answer produces. What architectural property made this possible, meaning a control plane rather than a person. Would the same technique work on another surface right now. What in the response was improvised, because every improvisation becomes a runbook entry or a pre-built capability.
Name a commander, open a timeline, and snapshot logs, assembled prompts, model and config versions and the tool registry as it stands, inside five minutes.
Disable tools, degrade the surface or pull it. Revoke every credential the system holds and any sharing its scope. Containment comes before scope, every time.
What could this reach, from the tool registry and credential scopes, in minutes. Brief legal on the envelope, because notification clocks start at discovery rather than confirmation.
Every review ends with a control added, a scope narrowed, a detection created, a runbook entry written or a decision formally made and tested. A write-up alone is not finished.
Declare and name a commander. Preserve evidence before changing anything. Contain by disabling tools, degrading the surface or pulling it. Revoke and rotate any credential the system holds. Establish the blast radius envelope from the tool registry rather than waiting for actual scope. Notify legal if the envelope includes personal or customer data. Then check other surfaces for recurrence.
From the retrieval logs, never the complaint. Search the corpus for instruction-shaped text near the affected retrievals, identify the ingestion path, then enumerate every session that retrieved the affected content. Then examine what those sessions did next: tool calls, fetches, and what was returned to whom.
Legal owns that decision and that channel, and triggers vary by jurisdiction, sector and contract. What the runbook does is get legal the inputs early: discovery time, data classes, record counts, affected parties and jurisdictions. Notification clocks in several statutes start at discovery rather than confirmation.
Stop it, with an emergency stop if you have one and credential revocation if you do not. Do not pause and observe, because an acting agent is an expanding incident. Revoke every credential it holds and any sharing its scope. Snapshot the action audit trail and tool configuration before a deploy overwrites them, then reconstruct forward from the trigger.
The model still responds but with tools disabled, retrieval disabled and a restricted capability set. Without it your only containment options are full disable or nothing, and that argument gets had under pressure with a product manager in the room. It costs roughly a day per surface to build and cannot be improvised.
Treat it as one until characterised. Pin to the previous version if pinning is available, and degrade the surface if it is not. Run your evaluation and security regression suites against the new version immediately rather than relying on release notes. Then identify which controls depended on the previous behaviour, because format changes break validators.
Drop your details and we'll send AI Security Operations Runbook straight to your inbox - no spam, unsubscribe anytime.
It costs about a day per surface and cannot be improvised mid-incident. Work through your containment options with our engineering leads. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers