IR Ranking Metrics

You cannot tune retrieval without an offline metric and a judged set of queries. Standard metrics:

  • Precision@k / recall@k: share of the top k that is relevant / share of all relevant documents found in the top k. Order-unaware.
  • MRR (mean reciprocal rank): average of 1/rank of the first relevant result; suits single-answer lookups.
  • MAP: mean over queries of average precision across relevant positions; binary relevance.
  • NDCG@k: graded relevance with a log position discount. DCG = sum(rel_i / log2(i+1)), normalised by the ideal ordering (nDCG = DCG/IDCG) so scores run 0 to 1 and compare across queries. The usual headline metric for ranked lists (BEIR reports nDCG@10).

Judgments

Relevance labels can be human, LLM-judged, or implicit (clicks, dwell). Offline metrics and online outcomes (task success) can diverge, so validate important changes with live experiments.

For RAG

Retrieval metrics cover only the retriever. Recall@k matters most for the generator (the answer needs the evidence somewhere in context); precision and ordering matter for reranking and context budget. Ragas offers LLM-based context precision and context recall that work without gold rankings; see rag-evaluation.

rerankers-and-cross-encoders · hybrid-search-and-rank-fusion · _rag-relevance-moc

Sources