Spots

How to Evaluate RAG Pipeline Quality: Metrics and Test Harness

Most teams building RAG systems spend 90% of their time on retrieval and generation, and 10% on evaluation. That ratio is backwards. Without a rigorous test harness, you're shipping a black box — and you'll only find out it's broken when a user does. What You're Actually Measuring

A RAG pipeline has two distinct failure modes

A RAG pipeline has two distinct failure modes: retrieval failures (the right context isn't fetched) and generation failures (the model hallucinates or ignores the retrieved context). Your evaluation framework needs to catch both independently. The four metrics that matter in practice: Context Recall: Did the retriever find the documents needed to answer the question? Context Precision: Of what was retrieved, how much was actually relevant? Faithfulness: Does the generated answer stick to the retrieved context, or does it introduce invented facts? Answer Relevance: Does the answer actually address the question asked?

You can measure these with a reference dataset

You can measure these with a reference dataset (ground-truth question/answer/context triplets) or, more scalably, with a language model acting as a judge. Both approaches are worth understanding. Building a Reference Dataset

The foundation of any test harness is a

The foundation of any test harness is a dataset of question/answer/context triplets. For each item you need: the question, the correct answer, and the document chunks that should be retrieved.

Build this dataset by exporting real user queries

Build this dataset by exporting real user queries paired with your best-known answers and the source chunks. Even 50 well-curated samples will catch most regressions. Seed it with edge cases: short questions, ambiguous phrasings, questions that span multiple documents. Computing Faithfulness with an LLM Judge

Faithfulness is hard to measure with string matching

Faithfulness is hard to measure with string matching — you need semantic understanding. The standard approach is using a language model as a judge. The judge receives the retrieved context and the generated answer, then returns a score from 0 to 1.

A faithfulness score below 0.8 typically means the

A faithfulness score below 0.8 typically means the retriever is fetching irrelevant chunks and the generator is hallucinating from them — a retrieval bug that looks like a model bug until you measure it properly. Computing Context Recall with Token Overlap

For context recall — "did we retrieve what

For context recall — "did we retrieve what we needed?" — a simple token overlap heuristic works well as a fast baseline before reaching for a semantic similarity model:

This is fast, zero-cost, and good enough to

This is fast, zero-cost, and good enough to catch retrieval regressions in CI. For production monitoring, replace it with a semantic similarity model such as a bi-encoder from sentence-transformers — token overlap misses synonyms and paraphrases. The Test Harness: Putting It Together

The p10 faithfulness score — the 10th percentile

The p10 faithfulness score — the 10th percentile — matters more than the mean. Your pipeline's worst cases are what end up in user complaints.

News

How to Evaluate RAG Pipeline Quality: Metrics and Test Harness

Most teams building RAG systems spend 90% of their time on retrieval and generation, and 10% on evaluation.

@spots #dev
Source: Dev.to
See more like this