Intermediate to senior

Machine Learning Interview Prep

Fifteen chapters from the learning problem and bias-variance to trees, neural networks, transformers, recommenders and ML system design, with tested NumPy code and diagrams.

Chapter 14 of 15Deep learning and applications · LLMs: Fine-Tuning, RAG and Evaluation

LLMs: Pretraining, Fine-Tuning, RAG and Evaluation

Large language model questions now appear in most ML interviews. They are usually practical: how a model is trained, when to prompt, retrieve or fine-tune, how parameter-efficient methods work, how to evaluate outputs, and what goes wrong (hallucination, injection, cost). This chapter builds the mental model and the decision framework. Product-specific numbers change quickly, so treat any model name or limit as something to verify, and focus on the principles.

1. The training pipeline of a modern assistant

  1. Pretraining. Next-token prediction on trillions of tokens of text and code. This produces a base model with broad knowledge and skills, but it only continues text; it does not follow instructions well. It is the most expensive stage by far.
  2. Supervised fine-tuning (SFT) / instruction tuning. Train on curated (instruction, response) pairs so that the model follows instructions and adopts a format and tone.
  3. Preference optimisation. Align behaviour with human or AI preferences. RLHF trains a reward model from ranked responses and optimises the policy against it with reinforcement learning (often PPO). DPO and related methods optimise preferences directly with a classification-style loss, avoiding a separate reward model and RL loop. Variants use AI feedback or verifiable rewards (such as tests passing or answers checked automatically), which is especially effective for reasoning and code.
  4. Safety training and evaluation, then deployment with system prompts, filters and monitoring.

Scaling laws describe how loss falls predictably with model size, data and compute, with an optimal balance between parameters and training tokens. Many modern models are trained on more tokens per parameter than early scaling work suggested, because inference cost depends on model size and smaller, longer-trained models are cheaper to serve.

2. Prompt, retrieve, or fine-tune?

The central decision. Start with the cheapest approach and escalate only when evidence says so.

NeedBest first toolWhy
Different output format, tone or task framingPrompting (clear instructions, examples)instant, no training
Answers grounded in your documents, fresh or private knowledge, citationsRAGknowledge lives in an updatable index, not in weights
Consistent style, structured output, a narrow task at low cost and latency, domain jargonFine-tuningchanges behaviour, can shrink the needed model and prompt
Reliable multi-step actions with external systemsTool use / agentsthe model calls APIs, code, databases
New facts that change oftenRAG, not fine-tuningfine-tuning is a poor way to inject facts and goes stale
Reasoning difficultystronger model, more test-time compute, better prompting

A useful rule: fine-tuning teaches behaviour and style; retrieval supplies knowledge. They combine well: a fine-tuned model that reads retrieved context in a stable format.

3. Prompting techniques

  • Be explicit about the task, audience, format and constraints.
  • Few-shot examples show the pattern; choose diverse, representative ones.
  • Chain of thought: ask for reasoning steps, which can improve multi-step problems at the cost of tokens. Reasoning-trained models do this internally.
  • Structured output: request JSON matching a schema, and validate it. Constrained decoding enforces a grammar.
  • Role and context in a system prompt; separate trusted instructions from untrusted data.
  • Decompose large tasks into steps or chained calls.
  • Evaluate prompts on a test set, not on anecdotes.

4. Retrieval-augmented generation (RAG)

RAG retrieves relevant passages and puts them in the prompt so the model answers from them.

Pipeline: ingest (parse, clean, chunk) → embed chunks → index in a vector store → at query time retrieve top- → optionally rerank → generate with the retrieved context and instructions to cite and to say when the answer is not present.

Design levers:

  • Chunking: by structure (headings, paragraphs) with modest overlap; size trades precision (small) against context (large).
  • Embeddings: domain-appropriate models; the same model for documents and queries.
  • Hybrid search: combine keyword (BM25) and vector search, which catches exact terms such as part numbers that embeddings blur.
  • Reranking with a cross-encoder improves precision of the final few.
  • Query rewriting and metadata filters (date, product, permissions).
  • Access control: retrieve only what the user is allowed to see.
import numpy as np

def cosine_top_k(query, docs, k):
    q = query / np.linalg.norm(query)
    d = docs / np.linalg.norm(docs, axis=1, keepdims=True)
    sims = d @ q
    order = np.argsort(-sims)[:k]
    return order.tolist(), sims[order]

docs = np.array([[1.0, 0.0, 0.0], [0.9, 0.1, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]])
query = np.array([1.0, 0.05, 0.0])
idx, sims = cosine_top_k(query, docs, 2)
assert idx == [0, 1] and sims[0] > sims[1] > 0.9

# scaling a vector does not change its cosine similarity: the metric ignores magnitude
assert np.allclose(cosine_top_k(query * 7, docs * 3, 2)[1], sims)

Failure modes: retrieval returns the wrong passages (most common), the answer needs several documents, chunks lose context, the model ignores the context, or sources conflict. Evaluate retrieval and generation separately: recall@k and MRR of retrieval; faithfulness and answer relevance of generation.

5. Fine-tuning

Full fine-tuning

Update all weights. Strong but memory-hungry (weights, gradients and optimiser state can total many times the model size) and each task needs a full copy.

Parameter-efficient fine-tuning (PEFT)

Freeze the base model and train a small number of added parameters.

LoRA (low-rank adaptation): represent the weight update as a product of two small matrices, with , and rank . Only and train. The adapter is tiny, swappable per task, and can be merged into the base weights for serving.

import numpy as np

d, k, r = 4096, 4096, 8
full = d * k
lora = r * (d + k)
assert lora / full < 0.005                          # under 0.5 % of the parameters of one weight matrix
assert lora == 65536

# the update starts at zero because B is initialised to 0, so training begins from the base model
rng = np.random.default_rng(0)
A = rng.normal(0, 0.02, (r, 64)); B = np.zeros((64, r))
assert np.all(B @ A == 0)

QLoRA keeps the frozen base in 4-bit precision while training LoRA adapters in higher precision, letting large models be tuned on a single GPU. Other PEFT methods: prefix tuning, prompt tuning, adapters.

Data quality dominates

A few thousand high-quality, diverse, correctly formatted examples often beat a large noisy set. Check for leakage with evaluation data, duplicated or contradictory examples, and a distribution that matches real use. Fine-tuning can cause catastrophic forgetting of general abilities; keep a general-capability regression test.

6. Evaluating LLM systems

LLM outputs are open-ended, so evaluation is a design problem.

MethodUseCaveat
Exact match / unit teststasks with checkable answers (math, code, extraction)cheapest and most reliable when applicable
Reference-based metrics (BLEU, ROUGE, BERTScore)translation, summarisationweakly correlated with quality for open-ended text
Human evaluationgold standard for quality, tone, safetyslow, costly, needs clear rubrics and agreement checks
LLM-as-judgescalable grading against a rubric, or pairwise comparisonbiases: favours longer answers, its own style, first position; calibrate against human labels
Task-level business metricsresolution rate, edit distance, user acceptancethe real target
Benchmarksbroad capability trackingcontamination and gaming; do not choose a model from public benchmarks alone

Build an eval set from real queries, covering common cases, hard cases and known failures; version it; run it on every prompt, model or retrieval change; slice results by category. Add adversarial tests (prompt injection, jailbreaks, unsafe requests) and regression tests.

Hallucination (confident but unsupported statements) is reduced by grounding (RAG with citations), instructing the model to abstain when unsure, verifying claims against sources, constraining output, and using tools for facts and calculations. It cannot be driven to zero, so design for it: show sources, add human review for high stakes.

7. Cost, latency and serving

  • Cost scales with tokens in and out; long prompts and long contexts dominate. Shorten prompts, cache repeated prefixes, retrieve fewer and better chunks, and route easy requests to smaller models (model cascades and routers).
  • Latency has two parts: time to first token (prompt processing) and per-token generation. Streaming improves perceived speed.
  • Serving optimisations: KV cache, continuous batching, quantisation, speculative decoding (a small draft model proposes tokens, the large model verifies them), tensor/pipeline parallelism, distillation.
  • Context windows are finite and attention is quadratic in length, and models can under-use information buried in the middle of long contexts. More context is not always better.
# a rough cost model: price per million tokens, input and output priced separately
def request_cost(in_tokens, out_tokens, price_in_per_m, price_out_per_m):
    return in_tokens * price_in_per_m / 1e6 + out_tokens * price_out_per_m / 1e6

# illustrative prices only; check the current price list
base = request_cost(6000, 500, 3.0, 15.0)
trimmed = request_cost(1500, 500, 3.0, 15.0)        # retrieve fewer, better chunks
assert trimmed < base and abs(base - 0.0255) < 1e-9
assert abs(trimmed - 0.012) < 1e-9

8. Safety and security

  • Prompt injection: untrusted text (web pages, emails, documents) contains instructions that hijack the model. Treat all retrieved or user-supplied content as data, restrict tool permissions, require confirmation for consequential actions, and filter outputs. There is no complete defence, so limit what a compromised model can do.
  • Data leakage: do not place secrets in prompts; enforce document-level access control in retrieval.
  • Jailbreaks, toxic or biased output, copyright and privacy issues: use policy filters, red-teaming, logging and incident processes.
  • Over-reliance: communicate uncertainty and keep humans in the loop for high-stakes use.

9. Agents and tool use (brief)

An agent loops: the model decides on an action (call a tool, search, run code), the system executes it, and the result goes back to the model, until it answers. Strengths: handles multi-step tasks and fresh data. Risks: compounding errors, cost, latency, looping, and security (tool misuse). Mitigate with narrow tools, step limits, validation of tool arguments, logging and human approval for side effects. Start with the simplest workflow that works; many tasks need only a fixed chain of calls, not a free-roaming agent.

10. Common mistakes

  • Fine-tuning to add knowledge that a retrieval layer would serve better and keep current.
  • Judging a system from a handful of examples instead of an eval set.
  • Trusting an LLM judge without calibrating it against humans.
  • Stuffing the context with everything retrieved.
  • No plan for hallucination, injection or cost.
  • Choosing a model from public benchmarks alone.
  • Evaluating retrieval and generation together, so you cannot tell which broke.

11. Practice questions

  1. Describe the stages of training an assistant model: pretraining, SFT, preference optimisation.
  2. When would you use prompting, RAG or fine-tuning? Give an example of each.
  3. Explain LoRA. Why does it need so few trainable parameters?
  4. How is DPO different from RLHF?
  5. Design and evaluate a RAG system over company documents. What would you measure?
  6. What are the weaknesses of an LLM-as-judge and how do you mitigate them?
  7. How would you reduce hallucinations in a customer-support assistant?
  8. How does prompt injection work and how would you limit the damage?
  9. Your LLM feature costs too much. List five ways to reduce cost without hurting quality.
Header Logo