Retrieve broadly
The first stage should produce a candidate set with good recall using vector, keyword or hybrid retrieval.
Rerank narrowly
A cross-encoder or LLM-based relevance step can spend more compute on a much smaller candidate set.
Keep the source metadata
Reranking should preserve stable evidence IDs and provenance so selected passages remain auditable.
Measure incremental value
Reranking adds latency and cost, so evaluate whether it meaningfully improves top-k relevance for your real queries.