Fifty-nine per cent correct arrives at a steering group as the health of a retrieval system, and it is two measurements added together. Either the passage holding the answer never reached the model, or it reached the model and the model misused it, and one figure cannot say which. So teams tune the half they happen to be good at, for quarters at a time, while the defect sits in a chunker splitting tables down the middle. This instrument scores sixteen rows against one query path, end to end, and every row wants the measurement behind it.
Add the four domain subtotals and read the band against this one query path. It does not grade the team. It tells you whether the next change you make can be aimed at the stage that is genuinely failing, or whether you are about to toss a coin.
Seventy-six and above. Every change lands on the stage that is actually failing, because the attribution sheet says which one it was. Keep the labelled set growing out of live queries and rescore whenever a corpus joins the path. The risk at this level is quiet: a set marked eighteen months ago stops resembling what people now ask.
Fifty to seventy-five, and most working pipelines land here. The insurer's policy assistant scored 51 while its steering group read 59.0% and concluded the model was weak. Close C1 and C3, the labelled set and the failure attribution, before the next tuning quarter opens, since every other fix is a coin toss until those two are shut.
Under fifty. One figure is being reported for two systems with nothing on hand to split it, so nobody in the room can be argued out of a hunch and the loudest engineer decides what gets touched. Stop the prompt work. Mark two hundred queries against real traffic, score again in a fortnight, and let the two counts choose.
C1 carries 8 points and wants the people who answer these questions daily, not the engineers who built the index. Two claims handlers took three days over two hundred production queries.
C2 is worth 8 points and wants nDCG beside it, per corpus rather than pooled, at the production k rather than a flattering fifty. Both belong next to the answer score, never instead of it.
C3 carries 7 points. Never retrieved, retrieved and misused, or correct. The three counts must sum to the sample exactly, and the rule you used to decide each case goes on the sheet.
D1 is worth 7 points. Post-filtering ranks over the whole corpus and then drops what the caller may not see, which thins context silently and collapses recall for your least privileged users.
One retrieval architecture: one corpus set, one chunking strategy, one index, one query path, scored end to end. Where several corpora feed the same path, take the weakest reading rather than the mean, since one unfiltered index decides what the whole path can safely serve.
They fail for unrelated reasons and different people fix them. A passage that never reached the model is a chunker, index or filter problem. A passage that reached it and was misused is a prompt, model or grounding problem. One score cannot tell you which of the two you have bought.
Corpus hygiene and provenance takes 26, chunking and indexing strategy 22, retrieval quality measured on its own 30, and grounding, permissions and evaluation discipline 22, across sixteen rows. Rows are all or nothing, and a metric living in a notebook on somebody's laptop scores zero.
It is an ordinary one for a pipeline that works. The insurer's policy assistant served four corpora and reported 59.0% answer quality every week. What it lacked was a labelled set and an attribution rule, which is why two thirds of its failures sat before the model and nobody could see it.
Not on its own. Recall@10 moved from 0.720 to 0.910 in that rebuild while the generation failure rate held at 18.1% exactly, yet the raw count of generation failures rose from 26 to 33. Under one number that rise reads as a regression and the prompt takes the blame.
This one goes deep on a single query path. Benchmarking Production AI scores an estate and hands back a distribution rather than one figure, which is where you find out how many paths like this you are actually running. Score the estate first, then bring the worst path here.
Two hours with the chunker configuration, the index metadata and the query logs open, plus whoever owns the corpora. Labelling the sample is the real cost: two people, three days, two hundred queries. Give every failed row an owner and a date before anyone leaves.
Drop your details and we'll send Benchmarking Retrieval Architecture straight to your inbox - no spam, unsubscribe anytime.
Pick the corpus your users complain about rather than the one that demonstrates well, bring whoever configured the chunker, and work the sixteen rows with our engineering leads. We label a sample against your own traffic. SECTION 7 - FAQ - 5 to 8 questions
Book a retrieval scoring session