Rerankers & Cross-Encoders
A reranker is a second-stage model that re-scores the candidates returned by fast first-stage retrieval, improving precision at the top.
- Bi-encoder: encodes query and document independently into vectors; efficient over millions of documents; used for first-stage retrieval.
- Cross-encoder: processes the query-document pair together and outputs one relevance score (0 to 1 in Sentence Transformers). More accurate but too slow for whole corpora, so it reranks a shortlist (for example the top 100). Nogueira and Cho’s BERT re-ranker improved MRR@10 on MS MARCO by 27% relative over the prior state of the art (2019).
- Late interaction (ColBERT): per-token vectors with MaxSim scoring; document vectors are precomputable, sitting between the two. See exploring-colbert-with-ragatouille.
- Hosted rerankers: Cohere Rerank currently lists rerank-v4.0-pro and rerank-v4.0-fast, with v3.5 and earlier still documented. Open-weight options (BGE, Jina, Voyage) exist but were not re-verified this session.
- Learning-to-rank (pointwise, pairwise, listwise objectives) trains rankers on judgments or click data; fine-tuning with LoRA is possible (low-rank-adaptation).
Standard recipe: hybrid retrieve (hybrid-search-and-rank-fusion), rerank to 5-10, prove the gain with ir-ranking-metrics, watch the added latency.
Related
query-understanding · rag-evaluation · embedding-models · _rag-relevance-moc
Sources
- Sentence Transformers retrieve & rerank: https://www.sbert.net/examples/applications/retrieve_rerank/README.html (accessed 2026-09-30)
- Passage Re-ranking with BERT: https://arxiv.org/abs/1901.04085 (accessed 2026-09-30)
- Cohere Rerank docs: https://docs.cohere.com/docs/rerank (accessed 2026-09-30)
- ColBERT: https://arxiv.org/abs/2004.12832 (accessed 2026-09-30)