You are buying an index: a vector database you will operate yourself, a managed retrieval service, or a platform that wraps both and sells you the pipeline too. Every figure quoted at you in that purchase is a genuine measurement of somebody else's workload.
The three questions that usually decide whether the thing works for you are the ones least often answered on paper, and they are where the filter runs, how long a deletion takes to become true, and who enforces permissions at query time.
This is the question set and how to read what comes back.
The worrying reply is rarely a lie. It is usually the honest answer of a competent team whose product has not yet met a corpus like yours, which describes a fair amount of this market. These three are worth recognising on the call rather than in month four.
described as filtering
The engine picks its candidates by similarity and then discards whatever fails the predicate, so a selective filter quietly starves the result list. What you want instead is the predicate evaluated during index traversal, a documented switch to exact search when selectivity is high, and the cost of each path stated plainly.
Said without qualification, and without mentioning tombstones, replica lag or when space is actually reclaimed, that answer means nobody has measured it. A usable reply names a bound: tombstoned on write, invisible to every replica inside a stated number of seconds, reclaimed at compaction, with every other store that holds a derivative listed.
Store the permissions as payload metadata and filter on them yourself, with the semantics left to you. Combine that with post-filtering and you have a disclosure problem rather than a configuration one. A good vendor evaluates entitlements inside the retrieval boundary and will run a negative test suite to prove it.
Scanned files, near-duplicate revisions, tables, the multilingual tail and the broken encodings. A clean sample measures the sample, and you already know how that story finishes.
Choose a predicate matching roughly one document in ten thousand, then measure recall and p95 against it. That single test separates pre-filtering from post-filtering in an afternoon.
Delete a known identifier, poll every store on a loop and record the second each one stops returning it. Then restore an older snapshot and check whether the record walks back in.
Take out vectors, payloads, identifiers and index settings, and load them somewhere else while the relationship is friendly. An export nobody has run is a clause, not a capability.
Ask where metadata filters are applied. A team that answers pre-filtering, names the index traversal it happens in, and volunteers the point at which the planner switches to exact search has operated this at scale. A team that calls post-filtering simply filtering has probably never met a selective predicate.
Unknowable from two numbers, and the pairing is the problem. Those figures usually come from separate runs at opposite ends of the parameter range. Ask for a curve of recall against p95 latency across a sweep, at your index size, with the memory footprint drawn on the same chart.
An erasure request you cannot evidence is an erasure you did not perform. Caches, replicas, snapshots and reranker stores all hold derivatives, and a restore from a snapshot taken before the request can reinstate content you have already told somebody is gone.
Run a negative suite rather than a positive one. Create a user entitled to nothing, write the questions most likely to drag a forbidden document into context, and require zero results at the recall setting you intend to ship on. Positive tests prove retrieval works, not that it refuses.
More than the re-embedding compute, which is the part most quotes cover. The real bill is the dual-index window, the evaluation rerun on your own query set and the cutover risk. Ask whether named vectors let two models live in the index together, since that is what makes the migration reversible.
Partly, but it is aimed at a purchase. Benchmarking Retrieval Architecture is the instrument for scoring a stack you already operate and want to improve. Use this one when you are choosing between vendors or renewing, and that one when the decision is what to fix next.
Heads of data and CTOs, with whoever will operate the index in the room. Procurement can send the list, but the replies only mean something to someone who will live with the write path once the pilot is over and the corpus has grown.
Drop your details and we'll send Retrieval Architecture: What Buyers Should Ask straight to your inbox - no spam, unsubscribe anytime.
Bring us the shortlist and the corpus you would rather not show anyone. A two-week trial sprint settles where the filter runs and how fast a deletion travels, before the contract says it does not matter. SECTION 7 - FAQ - 5 to 8 questions
Talk to our engineers