Build a known-answer set
Create representative queries with expected source documents or passages. Without expected evidence, retrieval quality is difficult to measure.
Evaluate retrieval first
Measure whether the correct evidence appears in the candidate set and how highly it ranks.
Evaluate grounding second
Check whether claims are actually supported by retrieved evidence and whether citations point to the right sources.
Add operational metrics
Latency, token usage, vector-query time, failure rate and cost per answered query are part of production quality.