Logiciel Solutions Contact Us
Success Stories Tech News Investors Contact Us
whitepaper

Benchmarking Retrieval Architecture.

Fifty-nine per cent correct arrives at a steering group as the health of a retrieval system, and it is two measurements added together. Either the passage holding the answer never reached the model, or it reached the model and the model misused it, and one figure cannot say which. So teams tune the half they happen to be good at, for quarters at a time, while the defect sits in a chunker splitting tables down the middle. This instrument scores sixteen rows against one query path, end to end, and every row wants the measurement behind it.

In depth

One Score For Two Stages Is Why The Tuning Quarter Fixed Nothing.

01

A single answer score adds two failure modes together.

A six-point fall could mean the index missed a third of the relevant documents, or it could mean somebody edited the prompt last Tuesday, and the reported figure separates neither. Split it and the argument in the room ends, since the three counts have to reconcile against the sample exactly. Report recall@k next to the answer score rather than in place of it.

In shortReport recall@k next to the answer score rather than…
02

Weights follow what decides the diagnosis, not what took longest to build.

Corpus hygiene and provenance carries 26 points, chunking and indexing strategy 22, retrieval quality measured on its own 30, and grounding, permissions and evaluation discipline 22. The thirty points sitting on retrieval alone are there because a team with no labelled set cannot attribute a single failure to anything. Write down the corpora, the chunker configuration, the embedding model digest and the production k first, since half the failed rows announce themselves while you do it.

In shortWrite down the corpora, the chunker configuration, t…
03

Fifty-one out of a hundred is what a working pipeline reaches.

A policy assistant at a European insurer, four corpora, scored 51 here while reporting 59.0% answer quality upward. Two claims handlers marked two hundred production queries in three days, and 56 of those queries had never retrieved the answering passage at all, against 26 that retrieved it and then misused it. Section-heading chunking and a fused lexical index took recall@10 from 0.720 to 0.910 in eleven days, and answer quality reached 74.5% with the prompt and model untouched.

In short5% with the prompt and model untouched
04

A rising generation failure count can mean nothing went wrong.

That rebuild moved generation failures from 26 to 33 while the generation failure rate held exactly: 26 of 144 is 18.1%, and 33 of 182 is 18.1%. More queries simply reached a generator already performing as well as it had before. Under a single score that rise reads as a regression, and somebody loses the following month to the prompt.

In shortsomebody loses the following month to the prompt
The detail

What The Total Says About Where Your Next Change Should Land.

Add the four domain subtotals and read the band against this one query path. It does not grade the team. It tells you whether the next change you make can be aimed at the stage that is genuinely failing, or whether you are about to toss a coin.

Zone · 01

Tuning the right half

Seventy-six and above. Every change lands on the stage that is actually failing, because the attribution sheet says which one it was. Keep the labelled set growing out of live queries and rescore whenever a corpus joins the path. The risk at this level is quiet: a set marked eighteen months ago stops resembling what people now ask.

Zone · 02

Guessing which half

Fifty to seventy-five, and most working pipelines land here. The insurer's policy assistant scored 51 while its steering group read 59.0% and concluded the model was weak. Close C1 and C3, the labelled set and the failure attribution, before the next tuning quarter opens, since every other fix is a coin toss until those two are shut.

Zone · 03

Tuning blind

Under fifty. One figure is being reported for two systems with nothing on hand to split it, so nobody in the room can be argued out of a hunch and the loudest engineer decides what gets touched. Stop the prompt work. Mark two hundred queries against real traffic, score again in a fortnight, and let the two counts choose.

By the numbers

The figures that make it a board-level conversation.

56 of 200
labelled queries whose answering passage was never retrieved, inside a system reported as 59.0% correct
74.5%
answer quality after the chunker and the index changed, with prompt and model untouched
40%
of organisations control access to their AI models and data, which Domain D turns into scored rows
Inside the report

What you'll take away.

01

Step 1 - Mark two hundred real queries before you tune anything

C1 carries 8 points and wants the people who answer these questions daily, not the engineers who built the index. Two claims handlers took three days over two hundred production queries.

02

Step 2 - Publish recall@k at the k that actually reaches the model

C2 is worth 8 points and wants nDCG beside it, per corpus rather than pooled, at the production k rather than a flattering fifty. Both belong next to the answer score, never instead of it.

03

Step 3 - Sort every failed answer into one of three buckets

C3 carries 7 points. Never retrieved, retrieved and misused, or correct. The three counts must sum to the sample exactly, and the rule you used to decide each case goes on the sheet.

04

Step 4 - Put entitlements in the query rather than after the ranking

D1 is worth 7 points. Post-filtering ranks over the whole corpus and then drops what the caller may not see, which thins context silently and collapses recall for your least privileged users.

Questions

Frequently asked.

What is the unit this scorecard measures?

One retrieval architecture: one corpus set, one chunking strategy, one index, one query path, scored end to end. Where several corpora feed the same path, take the weakest reading rather than the mean, since one unfiltered index decides what the whole path can safely serve.

Why separate retrieval from generation at all?

They fail for unrelated reasons and different people fix them. A passage that never reached the model is a chunker, index or filter problem. A passage that reached it and was misused is a prompt, model or grounding problem. One score cannot tell you which of the two you have bought.

How are the hundred points distributed?

Corpus hygiene and provenance takes 26, chunking and indexing strategy 22, retrieval quality measured on its own 30, and grounding, permissions and evaluation discipline 22, across sixteen rows. Rows are all or nothing, and a metric living in a notebook on somebody's laptop scores zero.

Is 51 out of 100 a poor result?

It is an ordinary one for a pipeline that works. The insurer's policy assistant served four corpora and reported 59.0% answer quality every week. What it lacked was a labelled set and an attribution rule, which is why two thirds of its failures sat before the model and nobody could see it.

We already report answer quality weekly. Is that not enough?

Not on its own. Recall@10 moved from 0.720 to 0.910 in that rebuild while the generation failure rate held at 18.1% exactly, yet the raw count of generation failures rose from 26 to 33. Under one number that rise reads as a regression and the prompt takes the blame.

How does this sit alongside your other benchmarking resources?

This one goes deep on a single query path. Benchmarking Production AI scores an estate and hands back a distribution rather than one figure, which is where you find out how many paths like this you are actually running. Score the estate first, then bring the worst path here.

How long does scoring take, and who should be in the room?

Two hours with the chunker configuration, the index metadata and the query logs open, plus whoever owns the corpora. Labelling the sample is the real cost: two people, three days, two hundred queries. Give every failed row an owner and a date before anyone leaves.

Get the whitepaper

Have it emailed to you.

Drop your details and we'll send Benchmarking Retrieval Architecture straight to your inbox - no spam, unsubscribe anytime.

Download whitepaper
Next step

Bring two hundred queries and we will split the number with you.

Pick the corpus your users complain about rather than the one that demonstrates well, bring whoever configured the chunker, and work the sixteen rows with our engineering leads. We label a sample against your own traffic. SECTION 7 - FAQ - 5 to 8 questions

Book a retrieval scoring session