Create representative fixtures
Keep examples that cover normal cases, edge cases and previous failures.
Test structured invariants
Schemas, allowed labels, required citations and tool-call constraints can be asserted automatically.
Score semantic quality separately
Some outputs need rubric-based or human evaluation because exact string matching is too brittle.
Compare before shipping
Run the same evaluation set against the old and new prompt/model and review deltas rather than judging a handful of examples manually.