Grade by hand
A classifier tagged each support ticket "urgent" or not. The gold label is the human's call. Flip a prediction and watch every score move.
Confusion matrix
Scores
Why accuracy lies
A model of 20 tickets that always predicts "not urgent." Slide how many are truly urgent and watch accuracy stay high while recall sits at zero.
Precision here is 0 / 0 (the baseline never predicts "urgent," so there is nothing to be right about), and it is guarded to 0.00. When zero tickets are urgent, recall is 0 / 0 too, also guarded to 0.00. Same rule as the code: an empty denominator scores 0.00, never a crash.