Rerankers & Cross-Encoders

A reranker is a second-stage model that re-scores the candidates returned by fast first-stage retrieval, improving precision at the top.

  • Bi-encoder: encodes query and document independently into vectors; efficient over millions of documents; used for first-stage retrieval.
  • Cross-encoder: processes the query-document pair together and outputs one relevance score (0 to 1 in Sentence Transformers). More accurate but too slow for whole corpora, so it reranks a shortlist (for example the top 100). Nogueira and Cho’s BERT re-ranker improved MRR@10 on MS MARCO by 27% relative over the prior state of the art (2019).
  • Late interaction (ColBERT): per-token vectors with MaxSim scoring; document vectors are precomputable, sitting between the two. See exploring-colbert-with-ragatouille.
  • Hosted rerankers: Cohere Rerank currently lists rerank-v4.0-pro and rerank-v4.0-fast, with v3.5 and earlier still documented. Open-weight options (BGE, Jina, Voyage) exist but were not re-verified this session.
  • Learning-to-rank (pointwise, pairwise, listwise objectives) trains rankers on judgments or click data; fine-tuning with LoRA is possible (low-rank-adaptation).

Standard recipe: hybrid retrieve (hybrid-search-and-rank-fusion), rerank to 5-10, prove the gain with ir-ranking-metrics, watch the added latency.

query-understanding · rag-evaluation · embedding-models · _rag-relevance-moc

Sources