Build a known-answer set

Create representative queries with expected source documents or passages. Without expected evidence, retrieval quality is difficult to measure.

Evaluate retrieval first

Measure whether the correct evidence appears in the candidate set and how highly it ranks.

Evaluate grounding second

Check whether claims are actually supported by retrieved evidence and whether citations point to the right sources.

Add operational metrics

Latency, token usage, vector-query time, failure rate and cost per answered query are part of production quality.