classification-metrics-and-threshold
Judge a classifier by what its errors cost, and set the decision threshold. Use when someone reports accuracy, precision, recall, F1, ROC-AUC or PR-AUC and asks what counts as good; when a model looks accurate but misses the cases that matter; when classes are imbalanced, as in fraud, churn, loan default, defects or medical screening; when choosing a probability cutoff; when comparing two classifiers; when reading a confusion matrix; when a review team can handle only so many alerts a day; or when checking whether predicted probabilities are calibrated. Not for choosing the algorithm or designing cross-validation, and not for ranking metrics for recommender systems.
Pinned to revision a18d88341e79, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/classification-metrics-and-threshold/SKILL.md
- skills/classification-metrics-and-threshold/scripts/threshold_by_cost.py
Every link opens the file at its source, pinned to the revision this page describes.