Retrieve broadly

The first stage should produce a candidate set with good recall using vector, keyword or hybrid retrieval.

Rerank narrowly

A cross-encoder or LLM-based relevance step can spend more compute on a much smaller candidate set.

Keep the source metadata

Reranking should preserve stable evidence IDs and provenance so selected passages remain auditable.

Measure incremental value

Reranking adds latency and cost, so evaluate whether it meaningfully improves top-k relevance for your real queries.