Create representative fixtures

Keep examples that cover normal cases, edge cases and previous failures.

Test structured invariants

Schemas, allowed labels, required citations and tool-call constraints can be asserted automatically.

Score semantic quality separately

Some outputs need rubric-based or human evaluation because exact string matching is too brittle.

Compare before shipping

Run the same evaluation set against the old and new prompt/model and review deltas rather than judging a handful of examples manually.