Intermediate to senior

Machine Learning Interview Prep

Fifteen chapters from the learning problem and bias-variance to trees, neural networks, transformers, recommenders and ML system design, with tested NumPy code and diagrams.

Chapter 1 of 15Foundations · How ML Interviews Work and the Learning Problem

How ML Interviews Work and the Learning Problem

Machine learning interviews test four things, usually across several rounds: whether you understand why methods work (theory), whether you can implement the core ideas (coding), whether you can frame a business problem as an ML problem and ship it (design), and whether you can reason about data and mistakes (judgement). This chapter maps the rounds, then sets up the vocabulary that every later chapter relies on: what learning is, what a loss is, and how to talk about generalisation.

1. The rounds you will meet

RoundWhat it checksHow to prepare
ML fundamentalsBias-variance, regularisation, metrics, trees, linear models, optimisationChapters 2 to 9 in this track. Be able to derive, not just recite.
CodingImplement k-means, logistic regression, a decision-tree split, softmax, or a metric from scratch in NumPyPractise writing without a library.
Applied or case study"Predict churn", "detect fraud", "rank search results"Chapters 12 and 13: framing, data, metric, baseline, deployment.
ML system designWhole pipeline: data, features, training, serving, monitoringChapter 13.
Statistics and probabilityBayes, distributions, hypothesis tests, A/B testingChapter 10.
Deep learning / LLMBackpropagation, attention, fine-tuning, evaluationChapters 7, 11, 14.
BehaviouralProjects, trade-offs, failuresPrepare two projects in depth.

Research roles weigh theory and papers; applied roles weigh data work and shipping; ML-engineer roles weigh systems. Read the job post and shift your time accordingly.

2. What "learning" means

You have inputs and outputs generated by some unknown process. You pick a model family (a set of functions with parameters ), a loss measuring how bad a prediction is, and an optimiser that adjusts to reduce the average loss on training data. This is empirical risk minimisation:

What you actually care about is the loss on new data from the same process, the generalisation error. The gap between training loss and that new-data loss is the central worry of the field.

TaskOutputTypical loss
Regressiona numbersquared error, absolute error, Huber
Binary classificationa probabilitylog loss (cross-entropy)
Multiclass classificationa probability vectorcategorical cross-entropy
Rankingan orderingpairwise or listwise loss
Clustering, dimensionality reductionno labelsreconstruction error, within-cluster variance

Supervised, unsupervised, self-supervised, reinforcement

  • Supervised: labelled pairs . Most production ML.
  • Unsupervised: only . Find structure: clusters, low-dimensional representations, anomalies.
  • Self-supervised: labels come from the data itself, such as predicting the next word or a masked patch. This is how large language models are pretrained.
  • Reinforcement learning: an agent acts, receives rewards, and learns a policy. Rarer in interviews outside research and robotics, but know the vocabulary: state, action, reward, policy, value.

Parametric versus non-parametric

A linear model has a fixed number of parameters regardless of data size (parametric). A k-nearest-neighbours model or a kernel method keeps the data and its complexity grows with it (non-parametric). Trees sit in between: the structure is learned from the data.

Generative versus discriminative

A discriminative model learns directly (logistic regression, trees, most neural classifiers). A generative model learns or (naive Bayes, Gaussian mixtures, language models, diffusion) and can sample new data.

3. A minimal learner, end to end

The smallest complete example: fit a line by gradient descent on a squared-error loss, and compare with the exact solution.

import numpy as np

rng = np.random.default_rng(0)
x = rng.uniform(-1, 1, size=200)
y = 3.0 * x + 0.5 + rng.normal(0, 0.1, size=200)          # true slope 3, intercept 0.5

X = np.column_stack([x, np.ones_like(x)])                  # add a bias column
w = np.zeros(2)
lr = 0.1
for _ in range(500):
    grad = 2 / len(y) * X.T @ (X @ w - y)                  # gradient of mean squared error
    w -= lr * grad

exact = np.linalg.lstsq(X, y, rcond=None)[0]               # the closed-form least-squares answer
assert np.allclose(w, exact, atol=1e-3)
assert abs(w[0] - 3.0) < 0.1 and abs(w[1] - 0.5) < 0.1

Every method in this track is a variation on this loop: choose a function family, define a loss, follow its gradient or split the data greedily, and judge the result on data the model did not see.

4. Train, validation and test

You need three roles for data, not two.

SplitUsed forRule
TrainingFitting parametersThe model sees these.
ValidationChoosing hyperparameters and comparing modelsYou may look at it repeatedly.
TestOne final, honest estimateTouch it once, at the end.

Every time you tune against a set, you leak a little of it into your choices. That is why the test set must stay sealed. With little data, replace a fixed validation set with k-fold cross-validation.

For time series and any data with a temporal order, split by time (train on the past, validate on the future). A random split lets the model peek at the future and gives a flattering, false estimate. For grouped data (multiple rows per user or patient), split by group so the same entity is not in both train and test.

5. The vocabulary interviewers expect

  • Parameters are learned (weights). Hyperparameters are chosen by you (learning rate, tree depth, regularisation strength).
  • Loss function is optimised; metric is what the business reads. They can differ, for example log loss for training and precision at k for evaluation.
  • Overfitting: low training error, high validation error. Underfitting: both are high.
  • Inductive bias: the assumptions that make a model prefer some functions over others (linearity, locality, tree-like interactions, translation invariance).
  • No free lunch: with no assumptions about the problem, no learner beats another on average over all problems. Success comes from matching bias to structure.
  • IID assumption: training and test examples come from the same distribution, independently. Production often violates it (drift), which is why monitoring matters.

6. Baselines come first

Before any model, build a baseline: predict the mean, the majority class, last period's value, or a simple rule. A model that cannot beat a one-line baseline is not worth deploying. In an interview, naming the baseline first signals that you think about value, not novelty.

import numpy as np

y_true = np.array([0, 0, 0, 0, 0, 0, 0, 0, 1, 1])          # 20 % positive
majority_pred = np.zeros_like(y_true)
accuracy = (majority_pred == y_true).mean()
assert accuracy == 0.8                                       # an "80 % accurate" model that learned nothing

That 80 % is the majority-class baseline. A model with 85 % accuracy here is barely better than doing nothing, which is why the metrics chapter insists on precision, recall and PR curves for imbalanced data.

7. How to answer a theory question

A structure that works for almost any "explain X" question:

  1. What problem does it solve? (one sentence)
  2. The core idea, in plain words, then the equation.
  3. A small example or picture.
  4. When it works and when it fails (assumptions).
  5. A comparison with the obvious alternative.

For example, "explain regularisation": it limits model complexity so the model fits signal and not noise; add a penalty to the loss; it shrinks weights, reducing variance at the cost of some bias; it helps when features are many or collinear; choose by validation; L1 gives sparsity, L2 gives smooth shrinkage.

8. Common mistakes in interviews

  • Reciting definitions without intuition or an example.
  • Jumping to deep learning for tabular data. Gradient-boosted trees are a strong default, and saying so shows maturity.
  • Ignoring the metric. "Accuracy" for a rare-event problem is a red flag.
  • Tuning on the test set, or splitting randomly on time-ordered data.
  • Not asking about data: volume, labels, how they were created, latency requirements, cost of errors.
  • Overclaiming. Say "this usually helps because ..., but I would check on validation data".

9. Practice questions

  1. Explain empirical risk minimisation and why training loss is an optimistic estimate of test loss.
  2. Why do we need a separate validation and test set?
  3. When would you use a time-based split rather than a random one?
  4. Give an example where the loss you train on differs from the metric you report.
  5. What is inductive bias? Compare the biases of a linear model, a decision tree and a convolutional network.
  6. What does the "no free lunch" theorem imply in practice?
  7. Construct a baseline for predicting next-week sales per store and say how you would beat it.
Header Logo