RAG Evaluation
A RAG system can fail at retrieval (wrong context) or at generation (right context, unsupported answer). Evaluation should score the two separately.
Axes
- Context relevance: did retrieval fetch the right passages? Retrieval-metric territory (ir-ranking-metrics).
- Faithfulness / groundedness: is each claim in the answer supported by the retrieved context?
- Answer relevance: does the answer address the question?
Ragas
Ragas (Es, James, Espinosa-Anke, Schockaert; arXiv 2309.15217) is a reference-free framework: metrics need no human-annotated ground truth and use LLM calls. The docs list Faithfulness, Context Precision, Context Recall, Response Relevancy, Context Entities Recall and Noise Sensitivity, plus agent metrics (Topic Adherence, Tool Call Accuracy, Tool Call F1, Agent Goal Accuracy). Other harnesses exist (TruLens, DeepEval, ARES, promptfoo) but were not checked this session.
Practice
- Build a fixed eval set of questions with reference contexts or answers; synthetic generation is common, but spot-check it.
- Run it on every retriever, reranker or prompt change as a regression test.
- LLM judges are biased and noisy: calibrate against human labels on a sample.
Related
retrieval-augmented-generation-overview · rerankers-and-cross-encoders · query-understanding · _rag-relevance-moc
Sources
- Ragas paper: https://arxiv.org/abs/2309.15217 (accessed 2026-09-30)
- Ragas metrics: https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/ (accessed 2026-09-30)