Intermediate to senior

Machine Learning Interview Prep

Fifteen chapters from the learning problem and bias-variance to trees, neural networks, transformers, recommenders and ML system design, with tested NumPy code and diagrams.

Chapter 8 of 15Training and evaluation · Evaluation Metrics and Imbalanced Data

Evaluation Metrics and Imbalanced Data

Picking and defending a metric is half of any applied ML interview. The right metric depends on the cost of each kind of mistake and on how the output will be used. This chapter covers classification, regression, ranking and probabilistic metrics, then the special problems of imbalanced data and thresholds.

1. The confusion matrix

For a binary classifier, every prediction lands in one of four cells.

<!--fig:confusion-->
Predicted Actual Positive Negative Positive Negative TPfound a real case FNmissed a real case FPfalse alarm TNcorrectly left alone Precision = TP / (TP + FP)how clean are my alertsRecall = TP / (TP + FN)how many real cases I caught Figure 1. The confusion matrix and the two metrics read from its rows and columns.
Predicted positivePredicted negative
Actually positiveTrue positive (TP)False negative (FN)
Actually negativeFalse positive (FP)True negative (TN)

From these:

MetricFormulaQuestion it answers
AccuracyWhat fraction is right overall?
PrecisionOf the items I flagged, how many are real?
Recall (sensitivity, TPR)Of the real items, how many did I find?
Specificity (TNR)Of the negatives, how many did I correctly leave alone?
False positive rateHow many negatives did I wrongly flag?
F1Harmonic mean of precision and recall
F-betaWeights recall times as much as precision
import numpy as np

def confusion(y_true, y_pred):
    tp = int(((y_true == 1) & (y_pred == 1)).sum())
    fp = int(((y_true == 0) & (y_pred == 1)).sum())
    fn = int(((y_true == 1) & (y_pred == 0)).sum())
    tn = int(((y_true == 0) & (y_pred == 0)).sum())
    return tp, fp, fn, tn

y_true = np.array([1, 1, 1, 1, 0, 0, 0, 0, 0, 0])
y_pred = np.array([1, 1, 0, 0, 1, 0, 0, 0, 0, 0])
tp, fp, fn, tn = confusion(y_true, y_pred)
assert (tp, fp, fn, tn) == (2, 1, 2, 5)
precision, recall = tp / (tp + fp), tp / (tp + fn)
f1 = 2 * precision * recall / (precision + recall)
assert abs(precision - 2 / 3) < 1e-9 and recall == 0.5 and abs(f1 - 4 / 7) < 1e-9

Why the harmonic mean?

The harmonic mean punishes imbalance between the two: precision 1.0 and recall 0.01 gives an F1 near 0.02, not 0.5. You cannot hide a terrible recall behind a perfect precision.

Choosing between precision and recall

  • Recall-heavy: missing a positive is costly. Cancer screening, fraud alerts at first pass, safety defects.
  • Precision-heavy: a false alarm is costly. Spam filtering that deletes mail, automated account bans, recommendations that annoy.
  • Both matter: F1, or set a precision floor and maximise recall.

2. Why accuracy fails

With 1% positives, a model that always predicts negative scores 99% accuracy and finds nothing. Accuracy is acceptable only when classes are roughly balanced and errors cost the same.

import numpy as np

y = np.array([0] * 990 + [1] * 10)
always_negative = np.zeros_like(y)
assert (always_negative == y).mean() == 0.99
recall = ((always_negative == 1) & (y == 1)).sum() / (y == 1).sum()
assert recall == 0.0

3. Thresholds, ROC and PR curves

A classifier outputs a score; a threshold turns it into a decision. Moving the threshold trades precision against recall. Two curves summarise all thresholds.

  • ROC curve: TPR (recall) against FPR for every threshold. AUC-ROC is the probability that a random positive is scored above a random negative. 0.5 is chance, 1.0 is perfect.
  • Precision-recall curve: precision against recall. Average precision (area under it) is the more informative summary when positives are rare, because ROC's FPR denominator (all the negatives) is huge and hides many false positives.
import numpy as np

def auc_rank(y_true, scores):
    """AUC as P(score of a random positive > score of a random negative), ties count half."""
    pos, neg = scores[y_true == 1], scores[y_true == 0]
    greater = (pos[:, None] > neg[None, :]).sum()
    ties = (pos[:, None] == neg[None, :]).sum()
    return (greater + 0.5 * ties) / (len(pos) * len(neg))

y = np.array([1, 1, 0, 1, 0, 0])
s = np.array([0.9, 0.8, 0.7, 0.6, 0.4, 0.3])
assert abs(auc_rank(y, s) - 8 / 9) < 1e-12              # 8 of 9 positive/negative pairs are ordered correctly
assert auc_rank(y, np.zeros(6)) == 0.5                  # a constant score is chance level

AUC cares only about ordering, not calibration. Two models can have identical AUC while one outputs probabilities that are useless as probabilities.

4. Probabilistic metrics and calibration

  • Log loss punishes confident mistakes severely. It is a proper scoring rule: it is minimised only by the true probabilities.
  • Brier score is the mean squared error of the predicted probability. Also proper, and bounded.
  • Calibration: among all cases where the model says 0.7, about 70% should be positive. Check a reliability diagram, and fix with Platt scaling (a sigmoid fit) or isotonic regression. Calibration matters whenever you act on the probability itself, such as pricing, risk scoring or combining model outputs.
import numpy as np

def log_loss(y, p, eps=1e-15):
    p = np.clip(p, eps, 1 - eps)
    return -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))

y = np.array([1, 0, 1, 0])
good = np.array([0.9, 0.1, 0.8, 0.2])
overconfident_wrong = np.array([0.99, 0.01, 0.01, 0.99])
assert log_loss(y, good) < log_loss(y, np.full(4, 0.5)) < log_loss(y, overconfident_wrong)

5. Multiclass and multilabel

  • Macro average: compute the metric per class, then average. Treats every class equally, good when rare classes matter.
  • Micro average: pool all decisions, then compute. Dominated by frequent classes (for single-label problems micro-F1 equals accuracy).
  • Weighted average: per-class metric weighted by class frequency.
  • Top-k accuracy when several guesses are allowed.

6. Regression metrics

MetricFormulaProperty
MAESame units as ; robust to outliers
MSE / RMSEPenalises large errors; RMSE is in 's units
MAPERelative, but explodes near zero and is asymmetric
Fraction of variance explained
Quantile (pinball) lossasymmetric absolute errorTrains and scores quantile forecasts

MAE is the median-optimal loss, MSE the mean-optimal loss. If a large error is disproportionately bad, use RMSE; if outliers are noise, use MAE.

7. Ranking and recommendation metrics

When the output is an ordered list, order matters near the top.

  • Precision@k, Recall@k: relevance in the top .
  • MAP: mean average precision across queries.
  • MRR: mean of .
  • NDCG@k: gains discounted by position and normalised by the ideal ordering, supports graded relevance.
import numpy as np

def dcg(rels):
    rels = np.asarray(rels, dtype=float)
    return np.sum((2 ** rels - 1) / np.log2(np.arange(2, len(rels) + 2)))

def ndcg(rels):
    ideal = dcg(sorted(rels, reverse=True))
    return dcg(rels) / ideal if ideal > 0 else 0.0

assert ndcg([3, 2, 1, 0]) == 1.0                        # already in ideal order
assert ndcg([0, 1, 2, 3]) < ndcg([2, 3, 1, 0]) < 1.0    # pushing relevant items down the list lowers the score

8. Imbalanced data

Rare events (fraud, defects, churn, disease) are the usual case in business ML. A checklist:

  1. Choose the right metric: PR curve, average precision, recall at a fixed precision, or cost-based metrics. Not accuracy.
  2. Stratify splits so each fold contains positives.
  3. Class weights: weight the rare class more in the loss. Simple and often enough.
  4. Resampling: undersample the majority or oversample the minority (SMOTE synthesises minority points). Do it only inside training folds, never before splitting, and remember it distorts predicted probabilities, so recalibrate afterwards.
  5. Threshold tuning: pick the threshold from validation data using the real cost of each error.
  6. Collect more positives or use anomaly detection if labels are almost absent.

Cost-sensitive thresholding

If a missed fraud costs ₹5,000 and a false alarm costs ₹50, flag any case whose predicted probability exceeds the break-even . The best threshold is far below 0.5.

c_fp, c_fn = 50, 5000
threshold = c_fp / (c_fp + c_fn)
assert abs(threshold - 0.0099) < 1e-4

# Expected cost of flagging a case with probability p of being fraud: (1-p)*c_fp. Not flagging: p*c_fn.
p = 0.02
assert (1 - p) * c_fp < p * c_fn                         # flag at 2 % even though 98 % are innocent

9. Offline versus online metrics

Offline metrics (AUC, NDCG) are proxies. The truth is an online experiment: an A/B test on the business metric (revenue, retention, click-through, time saved). Always say how you would close the loop: shadow mode, a small rollout, guardrail metrics (latency, complaints), and a pre-registered success criterion. A model with a better offline score can lose online if the metric is misaligned or the feedback loop changes behaviour.

10. Common mistakes

  • Accuracy on imbalanced data.
  • Reporting AUC when the decision depends on precision at a threshold.
  • Oversampling before the train/test split, which leaks duplicates into validation.
  • Ignoring calibration when the probability is used directly.
  • MAPE with values near zero.
  • Optimising a metric the business does not care about, or one the model can game.
  • Comparing models with a single run and no variance estimate.

11. Practice questions

  1. Define precision, recall and F1 and give an example where you would pick each as the primary metric.
  2. Why is PR-AUC more informative than ROC-AUC on a 0.1% positive rate?
  3. Explain AUC as a probability. What does it not tell you?
  4. A fraud model has 99.5% accuracy. Is it good? What would you ask?
  5. How do you choose a decision threshold given asymmetric costs?
  6. How do you evaluate a ranking model? Define NDCG.
  7. Describe two ways to handle class imbalance and one pitfall of each.
  8. Offline AUC improved but the A/B test shows no lift. What could explain it?
Header Logo