Query Understanding

Retrieval quality is capped by the query. Transforming it before retrieval is often the cheapest gain.

  • Expansion with LLM-generated text: query2doc (Wang, Yang, Wei, EMNLP 2023) prompts an LLM for a pseudo-document and appends it to the query; it reported 3% to 15% gains for BM25 on MS-MARCO and TREC DL without retraining.
  • HyDE (Gao, Ma, Lin, Callan, 2022): an LLM writes a hypothetical answer document, an unsupervised encoder embeds it, and that embedding retrieves real documents. The encoder’s bottleneck filters out hallucinated details. Competitive with fine-tuned retrievers without labels.
  • Rewriting: turn conversational, coreference-laden turns into standalone search queries.
  • Decomposition: split multi-part questions into sub-queries; the step toward agentic, multi-hop retrieval (see retrieval-augmented-generation-overview and _agentic-moc).
  • Routing and filters: classify intent to pick an index or lexical vs dense mode, and extract metadata filters (dates, entities).

Each rewrite costs an LLM call and can drift from user intent, so evaluate it against a baseline (ir-ranking-metrics, rag-evaluation).

hybrid-search-and-rank-fusion · rerankers-and-cross-encoders · _rag-relevance-moc

Sources