Observability, Security and Finishing Strong
The last minutes of a system design interview are where seniority shows. Anyone can draw the happy path. A senior candidate says how they would know the system is healthy, how it is protected, how it is deployed safely and what they would do next. This chapter covers those topics, then gives you a rubric and a practice plan.
1. Observability
You cannot operate what you cannot see. Three kinds of signal:
- Metrics: numeric time series, cheap to store and alert on. Request rate, error rate and latency percentiles.
- Logs: discrete events with detail. Structured logs (key-value) beat free text, because they can be searched and aggregated.
- Traces: the path of one request across services, with timing per hop. A trace identifier is passed along every call, so you can find which service made a request slow.
What to measure. Two checklists interviewers recognise:
- The four golden signals: latency, traffic, errors and saturation.
- RED for services (rate, errors, duration) and USE for resources (utilisation, saturation, errors).
Use percentiles, not averages. An average hides the slow tail. If 1 percent of requests take ten seconds, the mean looks fine and one user in a hundred has a terrible experience. A request that fans out to many backends is dominated by the slowest of them, so tail latency matters more as systems grow.
Alerting. Alert on symptoms users feel (error rate above the SLO, latency above target), not on every internal cause. Alerts must be actionable, or they are ignored. Tie alerts to the error budget.
Dashboards and runbooks let an on-call engineer who did not build the system diagnose it at three in the morning.
2. Safe deployment
- Rolling deployment: replace instances a few at a time, so capacity stays up.
- Blue-green: run the new version beside the old and switch traffic, with instant rollback by switching back.
- Canary release: send a small share of traffic, say 1 percent, to the new version, compare its metrics with the old, and widen only if healthy.
- Feature flags decouple deploying code from enabling a feature, and give a kill switch.
Database changes are the risky part. Make schema changes backward compatible in steps: add the new column, deploy code that writes both, backfill, switch reads, then remove the old column later. Never deploy a code version that requires a schema the running version cannot use.
3. Security essentials
Cover these in a sentence or two each, unprompted, for any design:
- Authentication (who are you) and authorisation (what may you do). Use a standard protocol for login and tokens, short-lived access tokens, and check permissions on the server for every request, never only in the client.
- Transport encryption (TLS) everywhere, including between internal services.
- Encryption at rest for sensitive data, with keys held in a managed key service and rotated.
- Secrets such as passwords and API keys are never in code or logs. Store them in a secret manager.
- Input validation and parameterised queries against injection. Limit request sizes.
- Rate limiting and abuse controls at the edge.
- Least privilege: each service gets only the access it needs, and a compromised service cannot reach everything.
- Passwords are stored as salted, slow hashes, not encrypted and never in plain text.
- Audit logs for sensitive actions.
- Personal data: collect what you need, define retention, and support deletion requests. Know the privacy laws that apply to your users, and say that you would confirm them rather than quoting from memory.
Common web threats to name: cross-site scripting, cross-site request forgery, injection, broken access control (the most common real-world failure), and denial of service. For the last, mention a CDN or scrubbing layer, rate limits and autoscaling with sensible caps.
4. Multi-region designs
Going multi-region improves latency and survives a regional outage, and it is expensive in complexity. Decide what you need:
- Active-passive: one region serves, another stands by. Simple, with a failover time and possible data loss for asynchronous replication, which you quantify as RTO (how long to recover) and RPO (how much data you may lose).
- Active-active: several regions serve, which needs a plan for data: partition users by home region, or accept conflict resolution for multi-leader writes.
Mention that cross-region round trips cost tens to hundreds of milliseconds, so synchronous replication across regions makes every write slow, and most designs replicate asynchronously and accept a small RPO.
5. Cost
Great designs are cheap enough to run. Notice where the money goes: data transfer between zones and regions, storage of cold data on hot tiers, over-provisioned capacity, and chatty services. Mention tiered storage, autoscaling, caching to cut compute and egress, and measuring cost per request.
6. Closing the interview
Spend the last few minutes on:
- A recap of the design in two or three sentences.
- The weakest points: "the metadata database is my single point of failure until I add replicas and a tested failover".
- Scaling limits: what breaks at ten times the load and the first thing you would change.
- What you would build next: monitoring, load testing, a staged rollout.
Naming your design's flaws before the interviewer does is one of the strongest signals available.
7. How you are likely to be scored
| Dimension | What a strong answer shows |
|---|---|
| Requirements | Clarified scope, stated assumptions, separated functional from non-functional |
| Estimation | Used numbers to choose the design, not as decoration |
| Design | A coherent path for each core use case; components justified |
| Depth | Went deep on a component, including failure modes |
| Trade-offs | Named what each choice costs, and when you would choose otherwise |
| Communication | Structured, checked in with the interviewer, adapted to hints |
| Operability | Monitoring, deployment, security and failure handling considered |
Level expectations differ. A mid-level candidate is expected to produce a sound, working design with guidance. A senior candidate drives the conversation, anticipates failures and chooses boundaries wisely. A staff-level candidate also challenges the problem, weighs organisational and cost dimensions and simplifies. Calibrate how much you lead to the level you are interviewing for.
8. A practice plan
- Learn the toolbox (chapters 2 to 6): load balancing, caching, storage and sharding, consistency, queues and reliability. Be able to explain each in two minutes.
- Practise the framework (chapter 1) on a timer until the sequence is automatic.
- Work the case studies here, then new problems: a rate limiter, a web crawler, a ride-hailing dispatch system, a payment system, a metrics platform, a search autocomplete service, a distributed cache. For each, write the estimate and the trade-offs on paper.
- Practise aloud, with a friend or tutor, and ask them to interrupt. Speaking a design is a different skill from thinking one.
- Review real systems. Engineering blogs and papers from large companies describe actual designs and the problems that forced them. Treat them as case studies of trade-offs, not as templates to copy.
- Keep a mistakes log. After each mock, write what you skipped, such as the estimate, a failure mode or the cost, and put it at the top of the next attempt.
Potential deep dives
Deep dive 1: How do you know the system is healthy?
The challenge. A system can be "up" and failing users.
Weak: monitor CPU and memory. Resource metrics say little about what users experience.
Solid: monitor the golden signals and alert on symptoms. Latency percentiles, traffic, errors and saturation, with alerts tied to user-visible symptoms and a service objective.
Excellent: objectives, burn rates and tracing. Define service level objectives from user needs and an error budget. Alert on the burn rate of the budget, so a fast burn pages immediately and a slow burn creates a ticket. Use distributed tracing with a correlation identifier to find where time goes across services, and structured logs for detail. Keep dashboards and runbooks for the on-call engineer, and review incidents to feed improvements back.
Deep dive 2: How do you deploy a risky change safely?
Weak: deploy to everyone at once. A bug reaches all users immediately.
Solid: a rolling deployment or blue-green. Replace instances gradually, or switch traffic between two environments with instant rollback.
Excellent: canary with automated judgement and flags. Send a small share of traffic to the new version, compare its error rate and latency with the old automatically, widen in steps and roll back on regression. Decouple deployment from release with feature flags, with a kill switch. Make database changes backward compatible in steps, so any version of the code can run against the current schema.
Deep dive 3: How do you secure the design?
Weak: add authentication and stop. Authentication says who. It does not say what they may do.
Solid: authenticate, authorise on the server, encrypt in transit and at rest. Short-lived tokens, permission checks on every request, TLS everywhere, secrets in a secret manager, input validation and rate limits.
Excellent: threat-model and apply least privilege. Identify the assets and the likely attackers. Give each service only the access it needs, so a compromise is contained. Treat broken access control (one user reading another's data) as the highest risk and test for it. Log security-relevant actions, protect keys and rotate them, minimise and retain personal data only as needed, and rehearse the response to a breach. Say that you would confirm the privacy and compliance rules that apply, and not quote them from memory.
Deep dive 4: How do you reason about cost?
Weak: ignore it. A design that is correct and unaffordable fails in practice.
Solid: notice the big cost drivers. Storage volume and tier, data transfer between zones and regions, compute headroom and managed service pricing.
Excellent: estimate and trade off. Estimate the monthly cost of the main components from the sizing numbers, find the dominant line, and reduce it with caching, tiered storage, compression, right-sizing and avoiding cross-region traffic. Compare building with buying, and say what the engineering time costs.
What is expected at each level
Mid-level. You mention monitoring and basic security, and can explain what you would alert on.
Senior. You propose concrete signals and objectives, a safe rollout strategy, security controls and a rollback plan, and you name the weakest part of your own design.
Staff. You tie operations to the organisation: ownership, on-call load, incident review, cost governance, and how the design will be maintained by people other than its author.
Interview questions and model answers
Q: How would you know this system is healthy in production? I would track the golden signals: latency percentiles, traffic, error rate and saturation, with alerts on user-visible symptoms tied to the SLO. Distributed tracing with a trace identifier through each call lets me find where time goes.
Q: How do you roll out a risky change? A canary to a small share of traffic, compared with the old version on error rate and latency, widened in steps, with a feature flag as a kill switch. Schema changes go in backward-compatible steps.
Q: What are the security basics for this design? Authenticate users with short-lived tokens, authorise on the server for every request, TLS everywhere, encrypt sensitive data at rest with managed keys, validate input, rate limit at the edge, and grant each service least privilege.
Q: What is the weakest part of your design? I name it honestly, for instance a primary database that is a single point of failure until replication and failover are in place, and I say how I would fix it and test it.
Common mistakes
- Finishing without any mention of monitoring or failure.
- Reporting averages when tail latency is what hurts.
- Treating security as an afterthought.
- Deploying schema changes that break the running version.
- Defending a design instead of discussing its limits.