Nobody has solved indirect prompt injection, so the question is not whether a vendor blocks it but what the agent can still reach once it has been persuaded. Most security due diligence never gets there. It collects certificates, asks about encryption at rest, and leaves the trust boundary undrawn. These seven themes put the architecture on the table instead, each paired with the fluent reply a well-briefed sales engineer gives when the thing underneath is a filter and a hope.
Fluency is not evidence. These three replies arrive in confident, well-rehearsed English, and each one tells you the architecture underneath is thinner than the sentence. Hearing any of them is a reason to stop and ask the follow-up in writing.
The most common reply, and the one that sounds most like an answer. Filters are a paraphrase, an encoding or a new language away from bypass, and no filter has ever been the thing that bounded damage. Ask what the model can still reach when the filter misses, and listen for whether they have ever measured that.
A deflection dressed as customer control. The agent holds whatever the integration was granted, which is usually one service connection with broad scope, so the answer describes who signed rather than what can be reached. Push until they name the credential behind each tool, its permitted actions, and the irreversible ones that stop at a human.
Offered as proof of isolation, it is a column in a table enforced by application code that one bug undoes. The separation worth buying is named per store, a namespace, a collection or a dedicated index, enforced below the application layer. Then ask about the cache, including the semantic one, and what happens when two tenants ask the same question.
Ask which component sees untrusted content and which one holds credentials. If the answer is the same component, the rest of the due diligence is decoration.
Every action the agent can take unattended, the credential behind it, which are read-only and which are irreversible. Then ask what one compromised session could chain together.
Scope, severities, what was fixed, what was accepted and why, and the retest date. OWASP Top 10 for LLM Applications is the minimum scope worth accepting.
A period in hours running from detection rather than confirmation, with interim updates. Add the rule that every tenant on a shared component hears about an injection, whether or not their data moved.
Ask them to show an injection payload that got through in their own testing and what it reached. A vendor with a real security programme names a bypass, a date, what it touched and what changed. One who says nothing has ever got through is telling you they have not looked.
No, and the gap is scope rather than rigour. Those exercises test the platform: network, access management, infrastructure. They rarely touch indirect injection through retrieved content, tool chaining or cross-tenant retrieval, which is where the expensive AI incidents start. Ask what categories were in scope and whether anyone attempted the corpora.
Treat it as a useful property and not as a control. Training reduces the hit rate; it does not bound what happens on the hits that land. The question that matters is what the model can still reach when it obeys the wrong instruction, and that is answered by tool scoping, not by weights.
More often than teams expect, and the refusal is itself information. Ask in writing, in these words, and give a deadline. A vendor with the artefacts sends them inside a week under NDA. A vendor without them sends a deck, and you have learned something worth more than the answer would have been.
Whoever can read an architecture diagram, plus the person who will own the incident. Procurement can send the list, but the follow-up questions are the valuable part and they need someone who knows what a pre-filter is. Budget an hour with the vendor's engineer rather than their account team.
This one points outward at a supplier. Benchmarking AI Security scores the systems you already run, across twenty controls and four domains. The AI Security Report prices what the incidents cost when the questions were never asked. Ask these before signature, score yourself after, and take the costs to the board.
Drop your details and we'll send AI Security: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.
A fortnight of free trial sprint produces code in your repo, a visible backlog, and an architecture note naming the trust boundary and the scope behind every tool credential. You keep all of it. SECTION 7 - FAQ - 5 to 8 questions
Book a due diligence review