Every AI vendor's security page says the right words. A questionnaire cannot distinguish a vendor who built these controls from one who knows what they are called, which is why most vendor security reviews end up measuring how good the vendor's security writer is. This scorecard puts five knockouts first, weights seven dimensions by real exposure rather than ease of assessment, and then hands you an architecture interview with the vendor's engineer, because that is the part of the process that actually separates built from described.
Three gaps recur across AI vendor assessments, and all three are invisible to a standard security questionnaire.
not one
Vendors answer isolation questions for the primary database and assume the rest is covered. AI systems leak between tenants in four other places: prompt and completion stores, vector indexes, response and context caches keyed in ways that can collide, and log stores a support engineer can read. A vendor who has not considered all four has not designed isolation.
For a hosted model provider you assess their infrastructure and isolation, because you own everything above the API. For AI-native software you assess the full stack, because they built the prompt assembly and tool layer you would otherwise own. For AI bolted onto an existing product, ask when the AI feature was security reviewed and by whom.
A vendor who will not share independent testing evidence, even under a non-disclosure agreement, is telling you something precise: either they have not been tested, or they have and would prefer you not read it. How routinely they handle that one request predicts the rest of the assessment better than any individual answer.
A hosted model provider, AI-native software, AI bolted onto an existing product and an agent platform need different questions and different weightings.
Tenant isolation across all four leak paths, encryption of prompt data, single sign-on with per-user access, testing evidence under NDA, and an agent permission model.
Ten questions, each paired with what a strong answer sounds like and what a weak one sounds like. Not a sales engineer, and record the refusal if you cannot get one.
Weak tenant isolation is a decline. Weak application security on an agent platform is a decline. No adversarial testing lowers confidence in every other dimension.
Start with tenant isolation across all four AI leak paths: prompts, embeddings, caches and logs, not just the database. Then whether prompt and completion data is encrypted in transit and at rest, whether they offer single sign-on with per-user access control, whether they will share independent testing evidence under NDA, and what the permission scoping model is for any agent features.
Substitute the architecture interview. Ask for an engineer rather than a sales engineer and listen for specificity: a strong answer names a layer, a constraint or a trade-off, while a weak one names a control. If they will not provide an engineer, score the application security dimension from the questionnaire alone and record the refusal as a finding.
What happens when your guardrail service is unavailable. It is narrow, factual, and cannot be answered from a security page. A vendor who has genuinely engineered their AI security has decided it and will say so in one sentence. One who assembled a stack will say they would have to check, and the answer is almost always fail-open.
Not the ones that matter here. It is meaningful evidence for infrastructure, access control and incident response, and you should ask for the full report rather than a bridge letter. It will not tell you whether retrieval enforces authorisation at query time, whether vector indexes are partitioned by tenant, or whether agent tool scopes are bounded.
Score that dimension no higher than one, and note explicitly that it lowers confidence in every other dimension rather than just its own, because nothing has been independently verified. Adversarial testing is the best available proxy for whether the other six dimensions are real or aspirational.
Especially then. That category has the most common gap: AI bolted onto a product whose security model predates it, often without the AI feature going through the same review as the rest. Ask directly when the AI feature was security reviewed and by whom, and weight tenant isolation, application security and adversarial testing heavily.
Drop your details and we'll send AI Security Vendor Scorecard straight to your inbox - no spam, unsubscribe anytime.
It cannot be answered from a security page. Bring us a vendor you are evaluating and we will sit in on the architecture interview. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers