Overview
Precision and recall are two classification metrics computed from the same confusion matrix but answering different questions about a model’s positive predictions. Precision asks how many predicted positives were correct, while recall asks how many actual positives were found. Optimizing one in isolation almost always trades off against the other.
Comparison Diagram
Comparison Table
| Aspect | Precision | Recall |
|---|---|---|
| Question answered | Of items predicted positive, how many actually are positive? | Of items that are actually positive, how many did the model find? |
| Formula | TP / (TP + FP) | TP / (TP + FN) |
| Denominator basis | Total predicted positive (TP + FP) | Total actual positive (TP + FN) |
| Error type penalized | False positives (false alarms) | False negatives (missed detections) |
| Increases when | Model makes fewer incorrect positive calls | Model catches more of the true positive cases |
| Threshold trade-off | Raising the decision threshold typically raises precision | Lowering the decision threshold typically raises recall |
| Failure mode at extreme | High precision, low recall: model is overly conservative and misses real cases | High recall, low precision: model is overly liberal and floods results with false alarms |
Key Differences
- Precision’s denominator is predicted positives; recall’s denominator is actual positives, so they measure against different totals
- Precision is hurt by false positives; recall is hurt by false negatives
- Adjusting the classification threshold pushes precision and recall in opposite directions
- Neither metric alone summarizes model quality, which is why the F1 score combines them
- A model with 100% recall can trivially predict everyone positive, and a model with 100% precision can trivially predict almost no one positive
When to Use Each
Precision
- Spam Filtering: Marking a real email as spam is costly, so you want confidence that flagged items truly are spam.
- Content Recommendation: Showing irrelevant recommendations degrades trust more than missing a few good ones.
- Automated Content Moderation: Wrongly banning legitimate users is more damaging than letting a few bad posts slip through.
Recall
- Disease Screening: Missing a true cancer case is far more dangerous than a false positive that gets ruled out later.
- Fraud Detection: Failing to catch fraudulent transactions costs more than flagging a few legitimate ones for review.
- Search and Retrieval Completeness: Legal or research search systems need to surface nearly every relevant document, even at the cost of some noise.