To evaluate RAG retrieval you need a labelled set of real queries paired with the documents that actually answer them, and then two numbers read at two different depths: recall at your candidate depth, which is the ceiling on what the system could possibly answer, and precision or nDCG at the depth the LLM actually reads, which is how much of that ceiling survives ranking. The metrics are the easy part and every vendor blog will hand you the formulas. The eval set is the work, and the one decision that makes it durable is labelling documents rather than chunks — so your labels survive the next time you re-index.
The one-line answer
Evaluate retrieval separately from generation, on real queries, with document-level labels, at the two k values your pipeline actually uses.
Everything that goes wrong with retrieval eval goes wrong upstream of the metric. Teams measure the wrong stage (end-to-end answer quality, which can’t tell you whether retrieval or the prompt failed), on the wrong queries (LLM-generated ones, which quietly test lexical echo), at the wrong depth (recall@5 on a pipeline that retrieves 50 and reranks), with labels that die the moment someone changes chunk size. Fix those four and a fifty-line scoring script is enough.
What you’re actually trying to decide
An eval exists to answer one question: if I ship this change, does the product get better? Not “what is our RAG score.” That framing matters because it tells you what to build and what to skip.
There are really only three decisions a retrieval eval has to serve:
- Should this change ship? Hybrid instead of pure vector, a reranker, a new chunking scheme, a different embedding model. This is a paired comparison on a fixed set — same queries, same labels, two configurations.
- Where is the system losing? Is the answer missing from the index entirely, or present and ranked too low? Those need opposite fixes, which is the whole argument in why your RAG retrieval is bad.
- Did we regress? A number in CI that fails a pull request when someone’s “harmless” filter change deletes 8% of recall.
Notice that none of those need an absolute score you can compare to another company’s. Retrieval numbers are not portable across corpora — a recall@50 of 0.82 on your contracts corpus tells you nothing about anyone else’s, and anyone quoting a benchmark number at you is comparing their corpus to your problem. Your eval is an instrument for your own decisions, and it only has to be consistent with itself.
Build the eval set first — it is most of the job
A mediocre metric on a good eval set beats a sophisticated metric on a bad one, every time. Budget accordingly: I expect to spend a day or two building the set and about an hour writing the scorer.
Get queries from real traffic. Production query logs first. If the system isn’t live yet, use the artefacts that already exist: support tickets, sales-engineering questions, the internal Slack channel where people ask the thing your bot is meant to answer, the search box on your docs site. Real queries are short, underspecified, full of internal jargon and misspellings, and often not questions at all — three keywords and a product code. Every synthetic query set I have seen is fluent, well-formed, and nothing like this.
Weight it toward failures, but don’t make it all failures. I aim for roughly half known-bad queries (the ones people complained about) and half ordinary traffic sampled at random. All-failures sets exaggerate every improvement and hide regressions on the queries that currently work — which is how you ship a change that fixes ten complaints and breaks a hundred silent successes.
Fifty to two hundred queries is enough to start. Below about fifty, a single query flipping moves your metric by more than most real changes do. Above a couple hundred, hand-labelling stalls and the set never gets finished. Start at fifty, ship the harness, grow the set every time a new failure gets reported — a live eval set that grows beats a perfect one that never launches.
Label documents, not chunks. This is the part I care most about. The obvious move is to record “for query Q, chunk 4f21a is the right answer.” Then you change your chunk size, re-index, and every chunk id in your eval set points at nothing. Your entire labelling investment evaporates on exactly the change you most needed to measure. Label the stable identifier instead — the document, or a section anchor that survives re-splitting — and score a retrieved chunk as a hit if its parent document is in the gold set. It costs you a little resolution and buys you an eval set that outlives your indexing decisions. That interaction is one reason chunking is the retrieval decision you can’t A/B cheaply; document-level labels are what make it A/B-able at all.
Graded labels where it’s cheap, binary where it isn’t. Binary (relevant / not) is fine for recall and MRR and takes a fraction of the time. You only need graded relevance (say 0–3) if you’re going to read nDCG seriously, because grading is what nDCG consumes. My default: binary labels, plus a 0–3 grade on a subset of thirty or so queries where ranking order genuinely matters.
On synthetic queries. LLM-generated question/chunk pairs are a legitimate cold-start tool when you have no traffic at all, and I’ve used them for exactly that. Know what you’re measuring, though. When you generate a question from a chunk, the question inherits that chunk’s vocabulary, so the retrieval task becomes “find the passage that shares my rare words” — a task lexical search wins trivially and that overstates every system’s performance. It also produces questions in the shape a model likes to ask, not the shape your users ask. Use them to get a harness running on day one, then replace them with real queries as fast as traffic allows, and never report the synthetic numbers as if they were your system’s.
Two numbers, two different jobs
The most common eval mistake I see is a single number at a single k that corresponds to no actual stage of the pipeline.
Modern retrieval is two-stage: retrieve broad, then rerank narrow, as covered in reranking in RAG. Those stages have different jobs and need different metrics.
Recall@N at your candidate depth — where N is however many candidates you pull before reranking, typically 50 or 100. For document labels, divide the number of distinct relevant documents retrieved by the number of relevant documents in the gold set. Missing evidence cannot be recovered by reranking those same candidates. A mean recall@50 of 0.70 means the average query recovered 70% of its labelled relevant documents; it does not mean 30% of queries were unanswerable. Read the zero-recall query count separately, and inspect whether the retrieved passages contain the required facts. This is the metric that helps you decide whether to work on the retriever — go hybrid, fix chunking, fix ingestion.
Precision@k and nDCG@k at your context depth — where k is how many chunks actually reach the LLM, typically 3 to 8. This is how much of the ceiling survives. If recall@50 is 0.92 and precision@5 is poor, your retriever is finding the evidence and your ranker is burying it: that’s a reranking or fusion-weighting problem, and it’s a much cheaper fix than re-indexing.
Reading both is the whole trick. One number tells you whether the system can find it; the other tells you whether the model ever sees it. Reporting recall@5 on a pipeline that retrieves 50 collapses those into one figure that can’t distinguish a retrieval failure from a ranking failure — the exact distinction you built the eval to make.
Two more metrics worth knowing and mostly not worth reporting:
- MRR (mean reciprocal rank) is the right metric when there is exactly one correct answer and its position matters — an internal lookup tool, a “which policy covers this” router. For multi-passage synthesis questions it throws away most of what you care about, since it only looks at the first hit.
- nDCG@k is the honest ranking metric when you have graded labels, because it rewards putting the most relevant thing first rather than merely somewhere in the window. Without graded labels it degenerates toward a fancier precision and I’d rather you just read precision.
A harness you can run today
Nothing here needs a framework. The whole scorer is this:
import math
from statistics import mean
def dcg(gains: list[float]) -> float:
return sum(g / math.log2(i + 2) for i, g in enumerate(gains))
def score_query(retrieved_doc_ids, gold, recall_at, context_k):
"""
retrieved_doc_ids: parent doc ids, in rank order, deduped, from the retriever.
gold: {doc_id: grade}. grade is 1 for binary labels, 0-3 if graded.
"""
gold_ids = {d for d, g in gold.items() if g > 0}
# Document coverage: a parent-id hit does not prove passage-level support.
top_n = retrieved_doc_ids[:recall_at]
recall = len(gold_ids & set(top_n)) / len(gold_ids) if gold_ids else 0.0
# Document-ranking proxy; this is not the final packed chunk window.
window = retrieved_doc_ids[:context_k]
precision = len([d for d in window if d in gold_ids]) / context_k
gains = [float(gold.get(d, 0)) for d in window]
ideal = sorted((float(g) for g in gold.values()), reverse=True)[:context_k]
ndcg = dcg(gains) / dcg(ideal) if any(ideal) else 0.0
rank = next((i + 1 for i, d in enumerate(retrieved_doc_ids) if d in gold_ids), None)
return {
"recall_at_n": recall,
"precision_at_k": precision,
"ndcg_at_k": ndcg,
"rr": 1.0 / rank if rank else 0.0,
"first_hit_rank": rank,
}
def evaluate(eval_set, retrieve, recall_at=50, context_k=5):
rows = []
for case in eval_set: # {"query": ..., "gold": {doc_id: grade}}
# dedupe chunks to parent docs, preserving rank order
seen, doc_ids = set(), []
for chunk in retrieve(case["query"], k=recall_at):
if chunk.doc_id not in seen:
seen.add(chunk.doc_id)
doc_ids.append(chunk.doc_id)
rows.append({"query": case["query"], **score_query(doc_ids, case["gold"], recall_at, context_k)})
summary = {m: mean(r[m] for r in rows) for m in ("recall_at_n", "precision_at_k", "ndcg_at_k", "rr")}
summary["misses"] = [r["query"] for r in rows if r["recall_at_n"] == 0]
return summary, rows
Two deliberate choices in there. The retrieved chunks are collapsed to parent documents before scoring, which is what makes document-level labels work. And evaluate returns the per-query rows, not just the means — because the means are the least useful thing it produces.
This harness measures document ranking, not faithfulness. Its context_k selects distinct parent documents after deduplication. Five distinct documents may require more than five chunks, and a hit can be the wrong passage from the right document. To evaluate the actual model input, also log the final passages after reranking, deduplication and token-budget truncation, and check their evidence coverage. Keep that context-level result separate from these document-level scores. Use a nonempty eval set with reviewed, nonempty gold sets for answerable queries; measure unanswerable queries through a separate refusal check.
Read the distribution, not the average
A mean recall of 0.80 is at least two very different systems: one that finds most of the evidence for nearly every query, and one that is perfect on 80% of queries and returns literally nothing for the other 20%. The second is a much worse product and a much easier fix, and the average hides which one you have.
What I actually look at, in order:
- The zero-recall list. Queries where nothing relevant came back at any depth. This is the highest-value list in the whole exercise; read the queries themselves and the failures usually cluster into two or three causes — an acronym the embeddings have never seen, a document type nobody ingested, a metadata filter excluding the answer.
- The count below a threshold, not the mean. “How many queries have recall@50 under 0.5” is a number that moves when the product improves and stays flat when you’ve shuffled the middle of the pack.
- Segments. Split by query type — keyword-ish lookups vs. natural-language questions, short vs. long, new documents vs. old. This is where the case for hybrid retrieval usually becomes undeniable: pure vector search craves the exact-token queries, which is the argument in hybrid search vs vector search, and a single mean smears that signal into nothing.
This is the same discipline as the difference between monitoring and alerting: an aggregate that always looks fine is not a measurement, it’s a comfort object. You want the number that changes when the thing you care about changes.
Turn the numbers into a ship gate
An eval that produces a report nobody acts on is theatre. Make it a gate:
- Paired comparison, never absolute. Run both configurations over the same eval set in the same run and diff per query. The interesting output is not “recall went from 0.78 to 0.81” — it’s the list of queries that flipped in each direction. A change that fixes twelve queries and breaks nine is a very different decision from one that fixes three and breaks none, and the means look identical.
- Ignore small deltas. With a hundred queries, a one-point move is noise. I don’t act on anything that doesn’t move several queries, and I look at which queries before I believe any of it.
- Regression-gate in CI. Freeze a small set — thirty queries is plenty — and fail the build if recall drops more than a couple of points. This catches the boring, expensive stuff: a filter change that silently excludes a document class, an ingestion job that stopped, an embedding-model version bump that shifted the vector space. Those are the failures that reach production, because nobody thinks they need testing.
- Re-label when you re-chunk. Document-level labels mostly survive; gold sets still drift as the corpus changes. Re-verify a sample every few months, and treat a query whose gold document was deleted as a data-quality bug, not a retrieval failure.
Where LLM-judge and end-to-end eval fit
Everything above measures retrieval. You still need to know whether the final answer was good, and that’s a separate instrument with separate failure modes.
Use an LLM judge for the things retrieval metrics can’t see: faithfulness (did the answer stay inside the retrieved evidence), citation correctness, refusal behaviour when the corpus genuinely doesn’t contain the answer. Those are real and worth measuring.
Treat judge output as an estimate to validate against human labels. Keep the judge model, prompt, claim-splitting rules and evaluation inputs fixed for comparisons; a change to any of those can change the score without a product improvement. Review disagreements and important failures by hand, and track judge cost and latency before adding it to every tuning run.
My split: retrieval metrics are the fast inner loop I run on every change, and judged end-to-end eval is the slower outer loop I run before a release. Debugging a bad answer with only an end-to-end score is the same mistake as debugging a slow request with only a total latency number — you know something is wrong and nothing about where.
Faithfulness versus retrieval recall: a worked example
Retrieval recall asks whether you found the relevant evidence. Faithfulness asks whether the answer’s claims are supported by the context actually supplied to the model. They have different denominators and can move in opposite directions.
For a claim-based faithfulness score, split the answer into factual claims and divide supported claims by total claims. This follows the Ragas faithfulness definition. Support means the context entails the claim; matching keywords or attaching a citation is not enough. It does not establish whether the source itself is correct or current.
The question and the gold evidence
Here is a fictional backup service, with two complete, one-sentence documents. These are hand-labelled teaching examples, not benchmark results or outputs from an automated judge.
Question: When does the service run backups, and how long does it retain them?
- Document A — schedule: “The service runs backups nightly at 02:00 UTC.”
- Document B — retention: “The service retains backups for 30 days.”
The gold document set is {A, B}. Each document supplies one required fact. Both runs below request two candidates, pass every returned document to the model unchanged, and use no further retrieval. That makes candidate recall@2 and final-context document recall equal in this example.
Run A: perfect faithfulness, incomplete answer
The retriever returns only A. The complete model context is the schedule sentence above. The answer is:
The service runs backups nightly at 02:00 UTC.
Count that schedule statement as one claim. A supports it, so faithfulness is 1/1 = 1.00. Only one of the two gold documents was retrieved, so document recall@2 is 1/2 = 0.50. The answer omits retention and fails to fully answer the question, despite its perfect faithfulness score.
The first repair is to find why B disappeared: ingestion, filtering, candidate selection or ranking. The answer should also make the missing retention information explicit rather than silently presenting a partial answer as complete. Inventing a retention period would not repair retrieval.
Run B: perfect retrieval recall, unsupported claim
The retriever returns A and B, and both sentences reach the model. The answer is:
The service runs backups nightly at 02:00 UTC. It retains them for 30 days. Customers can change the backup schedule.
Review the three claims against that exact context:
| Answer claim | Support in the model context |
|---|---|
| Backups run nightly at 02:00 UTC. | Supported by A. |
| Backups are retained for 30 days. | Supported by B. |
| Customers can change the backup schedule. | Unsupported: neither document says this. |
Now document recall@2 is 2/2 = 1.00, while faithfulness is 2/3 ≈ 0.67. The extra claim is unsupported, not proven false. In this run the required evidence is already present, so inspect generation instructions and unsupported-claim handling before changing the retriever. A corrected answer gives the two supported facts and removes the extra claim.
Keep the measurement boundaries explicit
Do not substitute generated-answer faithfulness for completeness: Run A shows why. Do not substitute a parent-document hit for passage coverage: the harness above can count A even if the retrieved chunk from A omits the schedule. Save the exact model context with each answer so you can tell those failures apart.
Metric names also need a definition. Ragas LLM-based context recall checks coverage of claims in a reference answer. Our document recall counts gold document IDs. Faithfulness checks claims in the generated answer. Record the unit, denominator, retrieval depth and pipeline stage with every score; the numbers are not interchangeable.
For these hand-labelled examples, an answer with no factual claims has no faithfulness denominator. Record it as not applicable and assess whether the refusal was appropriate. Do not silently award a perfect score to an empty answer.
When you don’t need this
I’d skip a formal retrieval eval when:
- The corpus is tiny. A few hundred documents where every query returns something a human can eyeball. Read the outputs; the ceremony costs more than it returns.
- You have an obvious, unfixed structural problem. If half your PDFs extracted as garbage or a document class was never ingested, measuring precisely how bad retrieval is delays a fix you already know you need. Fix the obvious thing, then measure.
- You’re pre-traffic and pre-users. Building an eval set out of imagined queries is a way of feeling rigorous while learning nothing. Ship something small, watch what people actually ask, build the set from that.
- Nobody will act on the result. If the answer to “recall is 0.6” is “we don’t have time to change anything,” you’ve bought a dashboard, not an instrument.
The moment any of those stops being true — real users, real corpus, someone about to spend a sprint on a change — the eval becomes the cheapest thing in the project. It’s the difference between knowing a change helped and believing it did.
FAQ
How many queries do I need in the eval set? Fifty to get signal on large changes, a couple of hundred to detect small ones. Below fifty, one query flipping swamps most real improvements. Weight toward known failures but keep a random sample of ordinary traffic, or you’ll never see the regressions.
Should I evaluate chunks or documents? Score at the document level. Chunk-level labels are more precise and become worthless the moment you change your splitter — which is exactly when you need the eval most. If you need chunk-level resolution for a specific investigation, do it as a one-off on a handful of queries rather than as the basis of your standing eval set.
Can I generate the eval set with an LLM? For a cold start, yes, and it beats having no harness. Just know that a question generated from a passage shares that passage’s vocabulary, so the numbers are optimistic and skewed toward lexical matching. Treat synthetic sets as scaffolding and replace them with real queries as soon as you have traffic.
Which single metric should I report if I only get one? Recall at your candidate depth. It’s the ceiling on everything downstream, and it’s the number that tells you whether to work on retrieval or ranking. Precision at context depth is the essential second number, not an optional one.
Do I need RAGAS or a dedicated eval framework? Not for retrieval metrics — that’s the fifty lines above, and writing it yourself means you know exactly what’s being measured at which k. Frameworks earn their keep on the generation side, where judge prompts, faithfulness scoring, and result tracking are genuinely fiddly to build well.
My recall is high but answers are still wrong. Now what? High document recall does not rule out a retrieval or context-assembly failure. Check the actual passages that reach the model: the right document can contribute the wrong chunk, or required evidence can be dropped during ranking or truncation. Once the required facts are present, investigate generation, faithfulness and answer correctness. The diagnostic order of operations walks that branch.
Can a RAG answer be faithful but incomplete or wrong? Yes. Run A above is fully supported but leaves half the question unanswered. An answer can also faithfully repeat an outdated or incorrect source. Measure evidence coverage, answer completeness and correctness separately from support against the supplied context.
Retrieval eval is unglamorous and it’s the thing that separates teams who improve their RAG system from teams who keep changing it. The pattern is always the same: real queries, document-level labels, two numbers at two depths, per-query output, and a gate that stops regressions before users find them.
If you’re weighing a retrieval change and can’t yet tell whether it helped, that’s the gap worth closing first — I write more about the whole pipeline in the hybrid search and RAG guides, and about the measurement discipline behind it in monitoring and operations.