Start from the error taxonomy
A missed high-risk case and an unnecessary review have different consequences. Product teams should define those failure types before selecting a metric.
Precision and recall answer different questions
Precision asks how trustworthy positive predictions are. Recall asks how much of the positive class is found. The right balance depends on the workflow.
Thresholds are product settings
The model score is not the final decision. Threshold selection determines workload, alert volume and risk exposure.
Evaluate the whole workflow
A useful evaluation also includes latency, data freshness, explanation quality, human review time and whether users can act on the output.