Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
framework

AI Security Maturity Assessment.

Your web application firewall does not read prompts. Your static analysis does not read system instructions. Your incident runbook has no entry for an agent that used its API token the wrong way. Application security rests on a boundary between instructions and data, and a language model dissolves that boundary by design, which means most of the programme you already have points at the wrong surface. The question is not whether your AI feature works. It is what happens when somebody attacks it, and that is measurable in ninety minutes.

In depth

The Guardrail Is Not The Boundary. The Credential Scope Is.

01

What happens by default: the AI feature is treated as another endpoint.

The existing programme is pointed at it, the provider's safety training is assumed to be a control, and a guardrail product is bought because it demos well. Tool scopes stay at whatever the framework defaulted to and prompt logs land in the observability lake that forty engineers can query. Then an incident reveals that nobody ever decided what happens when the guardrail times out, and the inherited answer is that it fails open.

In shortthe inherited answer is that it fails open
02

What good security does: it bounds the action surface before buying anything.

An agent's credential scope defines the worst case whether or not an attack was detected, so that comes first. Retrieved content is separated from instructions in prompt assembly, retrieval authorisation is enforced at query time rather than above it, and model output is schema-validated before anything downstream consumes it. Then indirect injection is tested through the real retrieval path, and detection efficacy is measured, because an untested detection is worse than none.

In shortbecause an untested detection is worse than none
The detail

Three Properties That Break Your Existing Controls.

AI security is not application security relabelled. Three properties of the architecture are why, and each one drives a domain in the assessment.

Zone · 01

Instructions and data share a channel

A retrieval system that ingests a customer-uploaded PDF is executing attacker-supplied text. Input validation cannot be a strict allowlist when the input is natural language and the instruction channel is the data channel. Injection is contained architecturally or it is not contained at all, and no product can impose that from outside your prompt assembly layer.

Zone · 02

Capability is granted at runtime

Blast radius is defined by configuration a product manager may have changed on Tuesday. An agent with a write-capable API token is a confused deputy with credentials. If you run agents with write, transact or deploy permissions and score below 14 on the agency domain, that is the highest-priority finding in the assessment regardless of your total.

Zone · 03

Behaviour is non-deterministic

The same input can produce a different output on the next call, and a provider can change model behaviour with no error and no deploy on your side. A control that passed once has not been shown to hold. That is why the evidence scale tops out at enforced-by-architecture rather than tested-successfully.

By the numbers

The figures that make it a board-level conversation.

1 in 4
malicious breaches are now AI-enabled, at an average cost of $6M
56%
increase in AI-driven attacks year on year, costing around $1M more per breach
38%
of organisations require IT approval before AI is deployed, down from 45%
Inside the report

What you'll take away.

01

Step 1 - Put the people who get paged in the room

A security or AppSec lead chairs, with the engineer who built the AI feature, the platform owner and whoever runs incident response.

02

Step 2 - Score eight domains against the evidence scale

Forty statements from attack surface and credentials through to detection and red teaming, each scored on demonstrated control rather than configured intent.

03

Step 3 - Apply the two override rules

Any domain below 8 out of 20 caps your band at Reactive. Write-capable agents with a weak agency domain outrank everything else on the page.

04

Step 4 - Work the remediation sequence in order

Eight actions ranked by exposure removed per unit of effort, starting with agent tool scopes because that is the only failure that destroys rather than leaks.

Questions

Frequently asked.

We already do penetration testing. Why is this different?

Scope. A conventional test covers the application surface around the model and will find real issues there. It will not usually test indirect prompt injection through your retrieval path, cross-tenant leakage in a vector index, or whether an agent's tool scopes exceed its documented purpose, because those need your prompt assembly and permission model rather than your HTTP surface.

Does this apply if we only call a hosted model API and train nothing?

Yes, and it is written primarily for that case. Almost nothing here concerns training. It concerns what you send to the model, what you let it read at runtime, what you let it do afterwards, and whether you would notice an attack. All of that is yours regardless of who trained the weights.

What if we have no agents, just a chat feature?

Score the agency domain as low-exposure and concentrate on input security, output handling and data leakage. The agent override will not apply, which usually moves the real finding to output handling. Model output reaching an interpreter, renderer or query builder without validation is the most common serious issue in chat-only deployments.

How much of this can we buy rather than build?

Less than vendors imply. Guardrails, observability and shadow AI discovery are genuinely purchasable and worth buying. Instruction and data separation, retrieval authorisation and tool permission design cannot be bought, because they are properties of how your system is assembled.

Can our engineers complete this, or do we need a specialist?

Your engineers, with a security lead chairing. It is written so the person who built the surface can answer it, because they are the only one who knows what the tool scopes actually are. If nobody in the room has read the production system prompt and looked at live tool permissions, you are not ready to score the middle domains honestly.

How does this relate to the governance assessment?

They share a scale and bands so results are comparable, and they deliberately overlap on inventory, vendor risk and incident response with different questions. Governance asks whether the process exists. Security asks whether it holds under adversarial pressure.

Get the framework

Have it emailed to you.

Drop your details and we'll send AI Security Maturity Assessment straight to your inbox - no spam, unsubscribe anytime.

Download framework
Next step

Bound the blast radius before you buy the guardrail.

Bring us your scored result and we will work through the highest-severity domain with your engineers. A working session, not a sales pitch. SECTION 7 - FAQ - 5 to 8 questions

Talk to our engineers