Overview

Precision and recall are two classification metrics computed from the same confusion matrix but answering different questions about a model’s positive predictions. Precision asks how many predicted positives were correct, while recall asks how many actual positives were found. Optimizing one in isolation almost always trades off against the other.

Comparison Diagram

Predicted PositiveActual PositiveFPTPFNwrong alarmsmissed casesPrecision = TP / (TP + FP)Recall = TP / (TP + FN)

Comparison Table

AspectPrecisionRecall
Question answeredOf items predicted positive, how many actually are positive?Of items that are actually positive, how many did the model find?
FormulaTP / (TP + FP)TP / (TP + FN)
Denominator basisTotal predicted positive (TP + FP)Total actual positive (TP + FN)
Error type penalizedFalse positives (false alarms)False negatives (missed detections)
Increases whenModel makes fewer incorrect positive callsModel catches more of the true positive cases
Threshold trade-offRaising the decision threshold typically raises precisionLowering the decision threshold typically raises recall
Failure mode at extremeHigh precision, low recall: model is overly conservative and misses real casesHigh recall, low precision: model is overly liberal and floods results with false alarms

Key Differences

  • Precision’s denominator is predicted positives; recall’s denominator is actual positives, so they measure against different totals
  • Precision is hurt by false positives; recall is hurt by false negatives
  • Adjusting the classification threshold pushes precision and recall in opposite directions
  • Neither metric alone summarizes model quality, which is why the F1 score combines them
  • A model with 100% recall can trivially predict everyone positive, and a model with 100% precision can trivially predict almost no one positive

When to Use Each

Precision

  • Spam Filtering: Marking a real email as spam is costly, so you want confidence that flagged items truly are spam.
  • Content Recommendation: Showing irrelevant recommendations degrades trust more than missing a few good ones.
  • Automated Content Moderation: Wrongly banning legitimate users is more damaging than letting a few bad posts slip through.

Recall

  • Disease Screening: Missing a true cancer case is far more dangerous than a false positive that gets ruled out later.
  • Fraud Detection: Failing to catch fraudulent transactions costs more than flagging a few legitimate ones for review.
  • Search and Retrieval Completeness: Legal or research search systems need to surface nearly every relevant document, even at the cost of some noise.