/knowledge/notes/thresholds-roc-auc
Concept note · Model Evaluation
Thresholds and ROC-AUC
Classifiers produce scores, and you choose a threshold. This note shows why that choice matters, what it controls and how to measure the trade-off.
- Studied
- Statistical Machine LearningCOMP90051
- When
- 2023 S1
- Applied in
- Human or Machine?
- Read / Refreshed
- ~5 min read2026-10-15
A classifier that returns a score or probability for each item has to answer one more question than you might think: not just "is this positive?", but "how confident do I have to be to say yes?" Lowering the threshold catches more positives but flags more false alarms. Raising it cuts the false alarms but misses positives. This note walks through the trade-off and how to measure it.
01
The idea
Most classifiers produce a score: logistic regression gives a probability, a linear model gives a value, a distance-based classifier gives a similarity. You choose a threshold and label anything above it as positive. That choice is not made once, in the model; it is a separate decision that depends on your goal.
If you are diagnosing a disease, missing a real case (a false negative) might be worse than testing someone who does not have it (a false positive). You might lower the threshold and catch more cases even though you will flag some healthy people. If you are approving a loan, the opposite holds: a false positive wastes money, so you raise the threshold to be stricter.
A confusion matrix records four counts: true positives (TP, you said yes and you were right), false positives (FP, you said yes and you were wrong), true negatives (TN, you said no and you were right) and false negatives (FN, you said no and you were wrong). Every metric is a ratio of these four.
02
The maths
True positive rate (sensitivity, recall) is the share of real positives that you caught:
False positive rate is the share of negatives that you incorrectly flagged:
Precision answers "of the things I said were positive, how many actually were?":
Recall is another word for TPR. F₁ is the harmonic mean of precision and recall, a single number that penalises if either is low:
A ROC curve plots TPR against FPR as you move the threshold. One corner is "flag everything as positive" (high TPR, high FPR). The opposite corner is "flag nothing" (low TPR, low FPR). A random guess traces the diagonal. A good classifier bows above the diagonal. The area under the curve (AUC) is a single number: 0.5 for random, 1 for perfect.
03
Try it
The widget below shows two overlapping distributions of scores, one for real positives and one for negatives. The slider sets your threshold. As you move it, watch how the confusion matrix, metrics and ROC curve change.
- Drag the threshold far left. TPR goes to 1 (you catch all positives) but FPR goes to 1 too (you flag all negatives by mistake).
- Drag it far right. TPR drops (you miss positives) and FPR stays near zero (you do not flag negatives).
- Find the elbow where TPR is high and FPR is still low. That is where the ROC curve bends away from the diagonal.
- Watch precision and recall diverge. Recall (TPR) improves as you lower the threshold, but precision (the share of your positive predictions that were right) gets worse because you flag more negatives by mistake.
04
Where I used it
05
Easy to get wrong
06
Sources
- Receiver Operating CharacteristicWikipediaClear overview of ROC, TPR, FPR and AUC with historical context from signal detection theory.
- Model evaluation: quantifying the quality of predictionsscikit-learn documentationComprehensive reference for all metrics, with code examples and when to use each.