The model is a set of weights. The index is a second copy of your documents, chunked and embedded, and the access controls the source system enforced rarely survive the trip. This is a lookup catalogue rather than an argument: twenty-eight numbered controls across seven families, each naming what it requires, the artefact a reviewer can hold, and the layer of the stack that owns it.
Two failures recur in every family. A control that sits in the interface instead of the retrieval call, and a control that sits in a notebook instead of a pipeline.
Seven families cover the catalogue, from corpus registration through to trace retention. These three carry most of what goes wrong once real traffic arrives, and each one is settled by an object somebody can open in front of you rather than by an assurance.
Caller identity resolves to an entitlement predicate evaluated inside the search call, never as a filter over results already returned. Entitlement changes reach index metadata within a stated interval, with revocations propagated ahead of grants. The artefact is a retrieval API contract that refuses any call arriving without a principal attached to it.
Freshness is published per corpus as the lag between source updated-at and indexed-at, reported at p95 rather than as a mean. Deleted identifiers go to a suppression list every build consults. A subject access request gets answered from the index itself, through a reverse lookup from a source record to the chunk ids derived from it.
Quality asserted by demonstration is what this family exists to prevent: a handful of questions work in a review, nothing is labelled, and a user reports the first regression. A labelled query set owned outside the build reports recall@k and nDCG per release, a gate blocks promotion below baseline, and every answer retains the index version that produced it.
Owner, source system, lawful basis, licence, sensitivity class and retention period, recorded before any data moves. Ingestion refuses a corpus that has no entry, and logs the rejection.
Parser version, chunker parameters, embedding model and dimension in a single file stored alongside the index. Regression test chunk boundaries so a parser upgrade cannot reshape the corpus silently.
Filtering results after retrieval means the model has read them already. Enumerate the service accounts reading the index too, since one non-human identity with corpus-wide read defeats the rest.
A similarity floor below which the retriever returns nothing, and a written contract for what the system says then. Vector search hands back its nearest neighbours whether or not they are relevant.
Ask whoever owns the retrieval service to show the API contract. If a search call can be made without a principal attached, entitlements are being applied somewhere above the retriever, which means the model has already read whatever matched before anything was hidden from the user.
Source document id, source system, document version, ingestion run id and ingested-at timestamp are what make the rest possible. Without them you cannot answer a subject access request from the index, cannot suppress a deleted record across a rebuild, and cannot trace a bad answer to the version of a document that caused it.
Keep a suppression list of deleted identifiers that the build manifest reads on every run, so a snapshot taken before the deletion cannot reinstate the content. Then test it: delete a known identifier, restore an older snapshot, rebuild, and confirm the record does not come back.
Quality you cannot measure is quality you cannot defend through a change. A labelled query set held outside the build gives you recall@k and nDCG against a recorded baseline, so an embedding swap or a chunker upgrade reports its own damage instead of waiting for a user to find it.
This one is the build. Retrieval Architecture Under Regulation covers which legal regime reaches which part of the stack and what evidence satisfies it, which is the argument. Read that for what is required of you, and this catalogue for the controls and artefacts that deliver it.
Data and platform engineers who own an index, and the technical leads reviewing their work. It assumes you have retrieval in production or close to it, and want the control numbers to cite in tickets rather than a case for why retrieval matters.
Permissions, in almost every case. A reproducibility gap costs you time and a measurement gap costs you confidence, but a retriever that returns what the caller may not read turns an engineering problem into a disclosure, and that one does not stay inside the team.
Drop your details and we'll send Retrieval Architecture: An Engineering Reference straight to your inbox - no spam, unsubscribe anytime.
Point a two-week trial sprint at one family, usually permissions or deletion, and find out which artefacts your retrieval stack can produce today. Working code in your repository, not a slide. SECTION 7 - FAQ - 5 to 8 questions
Book a retrieval review