“What is the difference between precision and recall?” is asked in nearly every machine learning interview, from new graduates to senior roles, because it reveals whether a candidate thinks about what a model’s errors cost. Reciting the formulas is the easy half. The harder half is the follow-up: “your fraud model has 0.84 precision and 0.63 recall, the business wants 90% of fraud caught, what do you do?” That question is about thresholds, validation data and trade-offs, and it is where most answers fall apart.
Before you start
You need to know what a binary classifier is and that most classifiers output a score or probability rather than a hard yes or no. The from-scratch code is plain Python; the rest uses scikit-learn 1.9 on Python 3.14, and printed outputs come from real runs. For heavily imbalanced data and PR-AUC, the companion guide on imbalanced classification goes further.
The short answer
From the confusion matrix, precision is TP / (TP + FP): of everything the model flagged, the share that was really positive. Recall is TP / (TP + FN): of all real positives, the share the model caught. F1 is their harmonic mean, 2PR / (P + R), so it is only high when both are. Raising the decision threshold usually increases precision and lowers recall. You optimise recall when misses are expensive (disease screening, fraud review queues) and precision when false alarms are expensive (blocking payments, auto-deleting content), and you pick the threshold that meets that requirement on validation data.
How it works
Every prediction falls into one of four cells. True positives (TP) are flagged and real. False positives (FP) are flagged but not real, also called type I errors. False negatives (FN) are real but missed, type II errors. True negatives (TN) are correctly left alone. Each metric is a ratio of these cells:
def confusion(y_true, y_pred):
tp = sum(t == 1 and p == 1 for t, p in zip(y_true, y_pred))
fp = sum(t == 0 and p == 1 for t, p in zip(y_true, y_pred))
fn = sum(t == 1 and p == 0 for t, p in zip(y_true, y_pred))
tn = sum(t == 0 and p == 0 for t, p in zip(y_true, y_pred))
return tp, fp, fn, tn
def prf(y_true, y_pred):
tp, fp, fn, _ = confusion(y_true, y_pred)
precision = tp / (tp + fp) if tp + fp else 0.0
recall = tp / (tp + fn) if tp + fn else 0.0
f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0
return precision, recall, f1
y_true = [1] * 5 + [0] * 15
y_pred = [1, 1, 1, 0, 0] + [1] + [0] * 14
print(confusion(y_true, y_pred)) # (3, 1, 2, 14)
print([round(v, 3) for v in prf(y_true, y_pred)]) # [0.75, 0.6, 0.667]Three of five positives were caught (recall 0.6), and three of the four alerts were right (precision 0.75). Accuracy is 17 / 20 = 0.85, which sounds better than either and hides that 40% of positives were missed.
Notice what each metric ignores. Precision never looks at false negatives, so a model that flags only its single most confident case can have perfect precision. Recall never looks at false positives, so flagging everything gives perfect recall. That is why they are reported together, and why F1 uses the harmonic mean: for P = 1.0 and R = 0.02 the arithmetic mean is a flattering 0.51, while F1 is 0.039. The harmonic mean is dominated by the smaller value.
The second idea is that precision and recall belong to a threshold. A classifier produces scores, and predict simply applies 0.5 to the probability. Lower the threshold and more cases are flagged: recall rises because you catch more positives, and precision usually falls because more negatives sneak in. The model is the same; only the operating point moves.
Step-by-step walkthrough
This walkthrough uses a dataset with about 11% positives and a gradient boosting model, then chooses an operating point properly. The data comes from make_classification(n_samples=20_000, n_features=20, n_informative=6, weights=[0.9], flip_y=0.02, class_sep=0.8, random_state=0).
Step 1: Split into train, validation and test
The threshold is a tuned parameter, so it needs its own validation data, separate from the test set used for the final report:
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split
X_tmp, X_test, y_tmp, y_test = train_test_split(X, y, test_size=0.25, stratify=y, random_state=0)
X_train, X_val, y_train, y_val = train_test_split(X_tmp, y_tmp, test_size=0.25,
stratify=y_tmp, random_state=0)
clf = HistGradientBoostingClassifier(random_state=0).fit(X_train, y_train)
p_val = clf.predict_proba(X_val)[:, 1]Stratifying keeps the positive rate the same in every split, so the validation set has enough positives (408 of 3,750) to give stable numbers.
Step 2: Sweep the threshold and watch the trade-off
from sklearn.metrics import precision_score, recall_score, f1_score, fbeta_score
for t in (0.2, 0.3, 0.5, 0.7):
pred = (p_val >= t).astype(int)
print(f"t={t}: precision={precision_score(y_val, pred):.3f} recall={recall_score(y_val, pred):.3f} "
f"f1={f1_score(y_val, pred):.3f} f2={fbeta_score(y_val, pred, beta=2):.3f}")
# t=0.2: precision=0.688 recall=0.828 f1=0.752 f2=0.796
# t=0.3: precision=0.758 recall=0.760 f1=0.759 f2=0.759
# t=0.5: precision=0.843 recall=0.630 f1=0.721 f2=0.663
# t=0.7: precision=0.905 recall=0.468 f1=0.617 f2=0.518The default 0.5 is not special: F1 peaks nearer 0.3, where precision and recall are balanced. F-beta generalises F1 by weighting recall beta times as heavily as precision; F2 prefers the low threshold, which suits problems where misses hurt more.
Step 3: Choose the threshold from the requirement
Suppose the requirement is “catch at least 90% of positives”. precision_recall_curve returns every achievable (precision, recall) pair with its threshold; take the highest threshold that still meets the recall target, since that keeps precision as high as possible:
from sklearn.metrics import precision_recall_curve
prec, rec, thr = precision_recall_curve(y_val, p_val)
ok = rec[:-1] >= 0.90 # the last point has no threshold
best = thr[ok].max()
print(round(best, 3)) # 0.07Step 4: Confirm on the test set
p_test = clf.predict_proba(X_test)[:, 1]
pred = (p_test >= best).astype(int)
print(f"precision={precision_score(y_test, pred):.3f} recall={recall_score(y_test, pred):.3f}")
# precision=0.502 recall=0.881Two lessons are visible. Meeting 90% recall costs a lot of precision: half of the alerts are now false. And the test recall is 0.881, slightly under the target, because a threshold tuned on one sample is a little optimistic on another. In production you leave a margin (target 0.92 on validation to deliver 0.90), monitor recall on labelled outcomes, and say the trade-off out loud to stakeholders.
Worked scenario
A trust and safety team ships a toxic-comment classifier with model.predict(), which uses a 0.5 threshold. Offline it reports precision 0.85 and recall 0.63, the same pattern as the test numbers above at 0.5 (precision 0.847, recall 0.631). Moderators review everything the model flags, and the policy requirement is that at least 90% of toxic comments reach a human.
The bug is not in the model; it is in treating 0.5 as part of it. Because flagged comments go to human review rather than being deleted automatically, a false positive costs a few seconds of moderator time while a false negative leaves abuse online. The fix is to replace predict with an explicit threshold chosen as in Steps 3 and 4, then size the review team for the new alert volume, since halving precision roughly doubles the number of comments humans must read per real case. If that volume is unaffordable, the conversation becomes “improve the model or relax the target”, which is a product decision informed by the curve, not a modelling detail hidden in a default.
Common mistake
- “Precision is the fraction of positives we found.” That is recall. A memory aid: precision is about the predictions, recall is about the real positives.
- Quoting precision and recall without the threshold or comparing two models at their default thresholds. Compare curves, or compare at the same operating requirement.
- Tuning the threshold on the test set, which turns the test set into validation data.
- Using F1 when costs are asymmetric. F1 weights both errors equally; use F-beta or an explicit cost.
- Averaging blindly in multiclass problems. With classes
ok(90 rows),spam(8) andabuse(2), a model that never predictsabusescores micro F1 0.96, weighted 0.949 and macro 0.612. Micro and weighted F1 hide the failing rare class; macro shows it.
Verify the behavior
Property-based checking is a quick way to prove a hand-written metric is correct, including the zero-division edge cases:
import random
from sklearn.metrics import precision_score, recall_score, f1_score
random.seed(0)
for _ in range(1000):
n = random.randint(1, 30)
t = [random.randint(0, 1) for _ in range(n)]
p = [random.randint(0, 1) for _ in range(n)]
ref = (precision_score(t, p, zero_division=0), recall_score(t, p, zero_division=0),
f1_score(t, p, zero_division=0))
assert all(abs(a - b) < 1e-12 for a, b in zip(prf(t, p), ref)), (t, p)
print("1000 random cases match scikit-learn")Follow-up questions
- Is precision monotonic in the threshold? No. Lowering the threshold adds both true and false positives; precision can rise briefly when the newly included cases happen to be mostly positive, which is why PR curves are jagged.
- What is specificity? TN / (TN + FP), the recall of the negative class; ROC curves plot recall against 1 - specificity.
- When is accuracy fine? When classes are roughly balanced and both errors cost about the same.
- How do you choose a threshold from costs? With calibrated probabilities, flag when
p > C_FP / (C_FP + C_FN); otherwise minimise total cost directly over thresholds on validation data.
Interview exercise
A cancer screening model flags 200 of 10,000 patients. Of the 120 patients who really have the disease, 90 are among those flagged. Compute precision, recall, F1 and accuracy, and say which number you would lead with to a clinician.
Answer and reasoning
TP = 90, FP = 200 - 90 = 110, FN = 120 - 90 = 30 and TN = 10,000 - 200 - 30 = 9,770. Precision is 90 / 200 = 0.45, recall is 90 / 120 = 0.75, F1 is 2 x 0.45 x 0.75 / 1.2 = 0.5625, and accuracy is (90 + 9,770) / 10,000 = 0.986. I would lead with recall, because in screening a missed cancer is the costly error: one in four patients with the disease is not flagged. Precision of 0.45 means more than half of flagged patients go on to a follow-up test they did not need, which is usually acceptable for a non-invasive follow-up. Accuracy is the least useful number here, since flagging nobody would score 0.988.
Continue learning
- Practise in the Machine Learning chapter and the machine learning MCQs.
- Go further with rare positives, PR-AUC and class weights in imbalanced classification.
- See where the probabilities come from in logistic regression interview questions.
- Reference: Google’s accuracy, precision and recall module and scikit-learn’s model evaluation guide.