Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

Retrieval Architecture: What Buyers Should Ask.

You are buying an index: a vector database you will operate yourself, a managed retrieval service, or a platform that wraps both and sells you the pipeline too. Every figure quoted at you in that purchase is a genuine measurement of somebody else's workload.

The three questions that usually decide whether the thing works for you are the ones least often answered on paper, and they are where the filter runs, how long a deletion takes to become true, and who enforces permissions at query time.

This is the question set and how to read what comes back.

In depth

The Numbers On The Data Sheet Are Real, And Not About You.

01

That benchmark ran on a corpus you will never have.

Public retrieval sets are tidy by construction, with clean text, one language, deduplication and a query set written to suit them, while yours holds scanned contracts whose character recognition layer invented words, four near-identical revisions of the same policy, tables that chunk into nonsense and a multilingual tail nobody has inventoried. A number from the tidy set describes the engine on its best day, measured by the party selling it. Ask which parameters produced that figure, and what those same parameters do to latency and memory.

In shortwhat those same parameters do to latency and memory
02

Recall and latency almost always come from two different runs.

Approximate nearest neighbour search is a dial rather than a property, so push it one way and the recall is wonderful at a cost nobody publishes, push it the other and the p95 looks excellent while the answers lose the one chunk that mattered. Filters are where that trade turns sharp. An engine that applies the predicate after the vector search has already chosen its candidates hands back four results when you asked for twenty, and does it exactly when the query mattered most.

In shortdoes it exactly when the query mattered most
03

Everybody demonstrates the read path and everybody lives on the write path.

A proof of concept loads a corpus once, runs a query set and produces a number, while production has documents changing hourly, people leaving groups, records that are deleted and have to stay deleted, and an embedding model that will eventually be deprecated. IBM found 97% of breached organisations had no AI-specific access controls, and prices a breach involving model inversion at $6.07M. A vendor who hands you an access list to store as metadata has moved an access-control duty into your application layer.

In shortA vendor who hands you an access list to store as me…
The detail

Three Replies That Should Stop The Conversation.

The worrying reply is rarely a lie. It is usually the honest answer of a competent team whose product has not yet met a corpus like yours, which describes a fair amount of this market. These three are worth recognising on the call rather than in month four.

Zone · 01

Post-filtering

described as filtering

The engine picks its candidates by similarity and then discards whatever fails the predicate, so a selective filter quietly starves the result list. What you want instead is the predicate evaluated during index traversal, a documented switch to exact search when selectivity is high, and the cost of each path stated plainly.

Zone · 02

Deletion happens immediately

Said without qualification, and without mentioning tombstones, replica lag or when space is actually reclaimed, that answer means nobody has measured it. A usable reply names a bound: tombstoned on write, invisible to every replica inside a stated number of seconds, reclaimed at compaction, with every other store that holds a derivative listed.

Zone · 03

Attach your access list

Store the permissions as payload metadata and filter on them yourself, with the semantics left to you. Combine that with post-filtering and you have a disclosure problem rather than a configuration one. A good vendor evaluates entitlements inside the retrieval boundary and will run a negative test suite to prove it.

By the numbers

The figures that make it a board-level conversation.

9
areas to settle in writing before a retrieval contract is signed
97%
of breached organisations ran no AI-specific access controls at all
$6.07M
average cost of a breach involving model inversion
Inside the report

What you'll take away.

01

Step 1 - Load the ugliest corpus you own

Scanned files, near-duplicate revisions, tables, the multilingual tail and the broken encodings. A clean sample measures the sample, and you already know how that story finishes.

02

Step 2 - Write the selective filter test before the first query

Choose a predicate matching roughly one document in ten thousand, then measure recall and p95 against it. That single test separates pre-filtering from post-filtering in an afternoon.

03

Step 3 - Put a stopwatch on deletion

Delete a known identifier, poll every store on a loop and record the second each one stops returning it. Then restore an older snapshot and check whether the record walks back in.

04

Step 4 - Perform the export during the pilot

Take out vectors, payloads, identifiers and index settings, and load them somewhere else while the relationship is friendly. An export nobody has run is a clause, not a capability.

Questions

Frequently asked.

What single question separates a serious vendor fastest?

Ask where metadata filters are applied. A team that answers pre-filtering, names the index traversal it happens in, and volunteers the point at which the planner switches to exact search has operated this at scale. A team that calls post-filtering simply filtering has probably never met a selective predicate.

They quoted recall of 0.95 and p95 latency of 20ms. Is that good?

Unknowable from two numbers, and the pairing is the problem. Those figures usually come from separate runs at opposite ends of the parameter range. Ask for a curve of recall against p95 latency across a sweep, at your index size, with the memory footprint drawn on the same chart.

Why does deletion timing matter so much commercially?

An erasure request you cannot evidence is an erasure you did not perform. Caches, replicas, snapshots and reranker stores all hold derivatives, and a restore from a snapshot taken before the request can reinstate content you have already told somebody is gone.

How do we test permissions without building the whole thing?

Run a negative suite rather than a positive one. Create a user entitled to nothing, write the questions most likely to drag a forbidden document into context, and require zero results at the recall setting you intend to ship on. Positive tests prove retrieval works, not that it refuses.

What does an embedding model change actually cost us?

More than the re-embedding compute, which is the part most quotes cover. The real bill is the dual-index window, the evaluation rerun on your own query set and the cutover risk. Ask whether named vectors let two models live in the index together, since that is what makes the migration reversible.

We already run retrieval. Does this still apply?

Partly, but it is aimed at a purchase. Benchmarking Retrieval Architecture is the instrument for scoring a stack you already operate and want to improve. Use this one when you are choosing between vendors or renewing, and that one when the decision is what to fix next.

Who should own these questions internally?

Heads of data and CTOs, with whoever will operate the index in the room. Procurement can send the list, but the replies only mean something to someone who will live with the write path once the pilot is over and the corpus has grown.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Retrieval Architecture: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Everyone demos the read path, so buy the write path.

Bring us the shortlist and the corpus you would rather not show anyone. A two-week trial sprint settles where the filter runs and how fast a deletion travels, before the contract says it does not matter. SECTION 7 - FAQ - 5 to 8 questions

Talk to our engineers