A governance item passes when the artifact exists. A security item passes only when somebody has tried to defeat it and failed. That difference in the standard of proof is the whole design of this checklist: the verification column asks what you did to confirm the control holds, not what you configured. It opens with a twelve-minute attacker walkthrough, because running a checklist against a weak design just produces a long list of failures that all trace back to two or three architectural decisions, and the walkthrough finds those while they can still be changed.
Three rules govern how the checklist is run. Together they are why a passed gate here means something a signed questionnaire does not.
Exceed the rate limit in a test environment and confirm it stops. Query as one user for data only another can see. Submit a document containing markup and instructions and inspect what reaches the model. Force a malformed response and read the log. Trigger a simulated attack and confirm the page arrived. Configuration is an intention; a test is evidence.
A surface is any distinct path where model output reaches a user, a record or an action. A product assistant, a support triage pipeline, an internal copilot and each agent are four surfaces, not one system, and they fail differently. Running the checklist against your AI in aggregate produces answers that are true somewhere and false somewhere else.
Users find unanticipated input paths within a week and attackers shortly after, and alert fatigue established in week one is permanent. At day thirty, read agent action logs by hand, re-review tool scopes against what was actually invoked, and re-run one injection attempt to confirm the control survived a month of releases.
Sit with the engineer who built the surface and answer twelve questions out loud. Three or more uncomfortable answers means a design conversation, not a checklist.
Per-service credentials in a secrets manager, workload identity rather than shared accounts, end-user authorisation propagated, and rate and spend limits proven by exceeding them.
Instruction separation, provenance tagging, schema validation, contextual encoding, retrieval authorisation and tool scopes, each tested through the real path rather than a mock.
Detection that has never been exercised is a statement of intent. Run the simulation, record time to page, then check the six red lines are all clear before sign-off.
The walkthrough, then the action and tools gate in full, then three items: output schema validation, contextual encoding, and retrieval authorisation at query time. Those cover excessive agency, unsafe output handling and cross-tenant leakage, which are the three paths where a single failure is either irreversible or immediately reportable.
Through the real path, not a mock. Plant an instruction in a document your system will actually ingest and attempt exfiltration. Return an instruction from a tool call and observe behaviour. Submit content containing markup and inspect what reaches the model after sanitisation. Testing the model with jailbreak prompts tests the model; testing the path tests your system.
An agent with write, delete, send, deploy or transact scope and no approval gate or ceiling. Retrieval running as a service account that reads every tenant. Model output reaching an interpreter, shell or query builder with no validation. Provider credentials in source control. Generated code executing with production credentials or egress. And no rollback, or one never executed.
Yourselves, and you should, because it is written so the engineer who built the surface can verify it. An external engagement finds classes of issue your team will not reach unaided and is worth booking once to calibrate, ideally at the end of your architecture work. Then convert their findings into automated tests.
Partly. You cannot verify a vendor's internals, so the gates become questions rather than tests, which is what the AI Security Vendor Scorecard is for. Use this checklist for surfaces you built and the scorecard for surfaces you bought, and register both.
What happens when the guardrail service is unavailable. It is narrow, factual, and cannot be answered from a security page. A team that has genuinely engineered its AI security has decided this and will say so in one sentence. A team that assembled a stack will say they would have to check, and the answer is almost always fail-open.
Drop your details and we'll send AI Security Readiness Checklist straight to your inbox - no spam, unsubscribe anytime.
Those two cover the irreversible and the reportable. Talk the rest through with our engineering leads. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers