IR Ranking Metrics
You cannot tune retrieval without an offline metric and a judged set of queries. Standard metrics:
- Precision@k / recall@k: share of the top k that is relevant / share of all relevant documents found in the top k. Order-unaware.
- MRR (mean reciprocal rank): average of
1/rankof the first relevant result; suits single-answer lookups. - MAP: mean over queries of average precision across relevant positions; binary relevance.
- NDCG@k: graded relevance with a log position discount.
DCG = sum(rel_i / log2(i+1)), normalised by the ideal ordering (nDCG = DCG/IDCG) so scores run 0 to 1 and compare across queries. The usual headline metric for ranked lists (BEIR reports nDCG@10).
Judgments
Relevance labels can be human, LLM-judged, or implicit (clicks, dwell). Offline metrics and online outcomes (task success) can diverge, so validate important changes with live experiments.
For RAG
Retrieval metrics cover only the retriever. Recall@k matters most for the generator (the answer needs the evidence somewhere in context); precision and ordering matter for reranking and context budget. Ragas offers LLM-based context precision and context recall that work without gold rankings; see rag-evaluation.
Related
rerankers-and-cross-encoders · hybrid-search-and-rank-fusion · _rag-relevance-moc
Sources
- Discounted cumulative gain: https://en.wikipedia.org/wiki/Discounted_cumulative_gain (accessed 2026-09-30)
- BEIR: https://arxiv.org/abs/2104.08663 (accessed 2026-09-30)
- Ragas metrics: https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/ (accessed 2026-09-30)
- Definitions of precision, recall, MRR and MAP are standard IR textbook material, not individually re-verified this session.