Mid to Senior Engineer

System Design Interview Prep

A structured path from the interview framework through core concepts, key technologies and patterns to eighteen full problem breakdowns, each with diagrams and weak, solid and excellent answers to every deep dive.

Chapter 6 of 36Core concepts · Reliability: Rate Limiting, Retries and Resilience Patterns

Reliability: Rate Limiting, Retries and Resilience Patterns

A design is judged by how it behaves when things go wrong. This chapter covers the controls that keep one failure or one noisy client from taking down everything, and the arithmetic of availability.

1. Availability in numbers

Availability is the fraction of time a system works. Convert it to allowed downtime, because the numbers are memorable:

AvailabilityDowntime per yearPer month (approx.)
99%3.65 days7.3 hours
99.9%8.76 hours43.8 minutes
99.99%52.6 minutes4.4 minutes
99.999%5.26 minutes26 seconds

Serial dependencies multiply. If a request needs three services each at 99.9 percent, the combined availability is about . Every hard dependency in the critical path lowers the total.

Redundancy raises it. Two independent replicas, each 99 percent available, give if either alone can serve. The word "independent" carries the weight: replicas in one rack or sharing one deployment pipeline fail together.

Related terms:

  • SLI (indicator): a measurement, such as the fraction of requests served in under 300 ms.
  • SLO (objective): the target for the SLI, such as 99.9 percent over 30 days.
  • SLA (agreement): a contractual promise, usually with penalties, set looser than the internal SLO.
  • Error budget: the allowed failure, 0.1 percent for a 99.9 percent SLO. When it is spent, feature work yields to reliability work.

2. Timeouts

Every network call needs a timeout. Without one, a slow dependency holds a thread, which holds a connection, and the caller's resources drain until it fails too.

Guidelines:

  • Set timeouts from observed latency, for example somewhat above the 99th percentile.
  • Make the sum of downstream timeouts smaller than the upstream timeout, or the caller gives up while its callees keep working on abandoned requests.
  • Propagate a deadline through the call chain, so deeper calls know how much time is left.

3. Retries

Retries recover from transient failures, and badly designed retries cause outages.

  • Retry only idempotent operations, or ones protected by an idempotency key.
  • Use exponential backoff: wait 100 ms, 200 ms, 400 ms and so on, up to a cap.
  • Add jitter: randomise each delay, so that thousands of clients that failed together do not retry together and re-crash the service.
  • Limit attempts and add a retry budget, for example no more than 10 percent of traffic may be retries.
  • Do not retry at every layer. If three layers each retry three times, one failure becomes twenty-seven calls. Retry at one chosen layer.

A retry storm is the classic failure: a dependency slows down, callers retry, the extra load slows it further, and it never recovers.

4. Circuit breakers

A circuit breaker stops calling a dependency that is clearly failing, so that callers fail fast and the dependency gets room to recover.

States:

  • Closed: calls pass through, failures are counted.
  • Open: after a failure threshold, calls fail immediately without being sent, for a cooling period.
  • Half-open: after the period, a few trial calls go through. If they succeed the breaker closes; if not, it opens again.

Pair a breaker with a fallback: a cached value, a default, or a reduced feature, so the user sees degraded service instead of an error. This is graceful degradation: the product recommendations fail and the page still loads without them.

<!--fig:breaker-->
failures over threshold cooling period ends trials succeed: close a trial fails: open again Closed calls pass, failures counted Open fail fast, no calls sent Half-open a few trial calls Figure 1. A circuit breaker fails fast while a dependency is down, then probes carefully before trusting it again.

5. Bulkheads and load shedding

Bulkheads isolate resources so one failing part cannot consume them all, like watertight compartments in a ship. Give each dependency its own thread or connection pool, and separate the capacity for critical and non-critical traffic.

Load shedding means refusing some work deliberately to stay healthy under overload. Return a fast "busy" response to low-priority requests, so that high-priority requests still succeed. A system that tries to serve everything under overload serves nothing well.

6. Rate limiting

A rate limiter caps how much a client may do in a period. Reasons: protect against abuse and accidents, share capacity fairly, control cost, and keep a system inside what it can serve.

Where to put it. At the edge or API gateway for coarse limits, and inside services for finer ones. Decide what you limit by: user, API key, IP address, or endpoint. State that IP-based limits hurt users behind shared addresses.

Algorithms

Fixed window counter. Count requests in each fixed interval, such as each minute. It is simple, but a client can send a full quota at the end of one window and another at the start of the next, doubling the burst at the boundary.

Sliding window log. Store a timestamp per request and count those in the last interval. It is exact, and it costs memory in proportion to the request count.

Sliding window counter. Blend the current and previous fixed-window counts by overlap. It is a cheap approximation of the sliding window.

Token bucket. A bucket holds up to tokens and refills at tokens per second. Each request takes a token and is rejected if none are left. It allows bursts up to and enforces a long-run average of . It is the most commonly chosen algorithm.

Leaky bucket. Requests enter a queue that drains at a fixed rate. It smooths bursts into a steady outflow.

Doing it across many servers

A limiter that counts in each server's memory limits each server, not the whole fleet. For a global limit, keep counters in a shared fast store and update them atomically.

-- token bucket in a shared store (sketch)
tokens, last = get(key)
tokens = min(B, tokens + (now - last) * r)
if tokens >= 1: tokens -= 1; allow
else: reject
set(key, tokens, now)   -- must be atomic with the read

The read-modify-write must be atomic, or two servers can both spend the last token. Use an atomic script or a compare-and-set. Accept that a small overshoot is usually fine, and that the limiter's own failure should fail open (let traffic through) for most APIs and fail closed where abuse is dangerous.

When rejecting, return HTTP 429 with a header telling the client when to retry, so well-behaved clients back off.

7. Health checks and failover

A liveness check says the process is running, and a readiness check says it can take traffic. Confusing them leads to restarting an instance that was only waiting for a dependency. Failover across zones or regions needs capacity to spare: if you run two zones at 70 percent each and one fails, the survivor cannot absorb the load. Size so that the remaining capacity covers the loss.

Chaos testing, deliberately killing instances or injecting latency in a controlled way, shows whether these mechanisms work before a real outage does.

Potential deep dives

Deep dive 1: How do you prevent a retry storm?

The challenge. A dependency slows down, callers time out and retry, the extra load slows it further, and it never recovers.

Weak: retry immediately, a few times, at every layer. Three layers retrying three times turn one failure into twenty-seven calls, all at the same instant.

Solid: exponential backoff with jitter and a limit. Wait longer after each failure, randomise the delay, and stop after a few attempts. Retry only idempotent operations.

Excellent: budgets, one retry layer and a breaker. Retry at one chosen layer. Cap retries to a fraction of traffic (a retry budget) so that retries can never dominate. Propagate a deadline so work is abandoned when the caller has given up. Add a circuit breaker so that callers stop sending to a dependency that is clearly down, and a fallback so users see a degraded response, not an error. Add load shedding at the dependency so it protects itself.

Deep dive 2: What does a good rate limiter protect?

The challenge. Clients can overload the system by accident or on purpose.

Weak: a single global limit. One heavy client consumes the allowance, and everyone else is rejected.

Solid: per-client limits with a token bucket. Each client has a budget with a burst allowance, rejected with a retry hint when exhausted.

Excellent: layered and failure-aware. Combine per-IP limits at the edge, per-account limits after authentication and per-endpoint limits for expensive routes. Enforce them atomically in a shared store. Decide failure behaviour per rule: fail open for ordinary traffic, fail closed for login and other abuse-sensitive paths. Pair limits with load shedding that protects the service when it is overloaded by legitimate traffic.

Deep dive 3: How do you design for graceful degradation?

The challenge. A non-essential dependency fails. The core flow should keep working.

Weak: any dependency failure fails the request. A broken recommendations service takes down the product page.

Solid: timeouts and fallbacks. Time out quickly and fall back to a default, a cached value or an empty section.

Excellent: classify dependencies and design the degraded modes. List each dependency as critical or optional. For optional ones, define the degraded behaviour in advance (hide the section, show cached data, use a simpler algorithm). Isolate resources with bulkheads so a slow optional dependency cannot exhaust threads for the critical path. Test degraded modes regularly, because an untested fallback is likely to fail. Use feature flags to switch features off under stress.

Deep dive 4: How do you survive a zone or region failure?

The challenge. A whole failure domain disappears.

Weak: assume it will not happen. It does.

Solid: replicate across zones, with capacity to spare. Run in at least three zones, with enough headroom that losing one still handles the load, and a tested failover for the data tier.

Excellent: set objectives and test them. Define recovery time and recovery point objectives for each service, pick active-passive or active-active accordingly, and weigh the cost. Keep the control plane independent of the data plane, so you can fail over while the failing region is down. Run game days that disable a zone on purpose, measure what happens, and fix the gaps. Do the capacity arithmetic: two zones at 70 percent load cannot absorb a failed zone, while three zones at 60 percent can.

What is expected at each level

Mid-level. You know timeouts, retries and basic rate limiting, and can convert availability percentages to downtime.

Senior. You prevent retry storms, use circuit breakers and bulkheads, design layered rate limits, define fallbacks and size failover capacity.

Staff. You set reliability objectives from user needs, use error budgets to balance features and reliability, test failure in production-like conditions, and design the organisation's response process.

Interview questions and model answers

Q: Design a rate limiter. Clarify who is limited (user or key), the limit and burst, and whether it is per server or global. I choose a token bucket for burst-tolerance, keep the bucket per key in a shared fast store updated atomically, return 429 with a retry-after header, and decide failure behaviour: fail open for a normal API. I mention sharding the counters by key and the cost of a hot key.

Q: Why can retries make an outage worse? They multiply load on an already struggling dependency. I use exponential backoff with jitter, a retry budget, retry at one layer only, and a circuit breaker so callers stop hammering a service that is down.

Q: What does 99.9 percent availability allow? About 8.8 hours a year, or 43 minutes a month. If my request depends on three services at that level in series, the combined figure is nearer 99.7 percent, so I limit hard dependencies in the critical path.

Q: How do you survive a dependency outage? Timeouts, a circuit breaker, a fallback such as cached or default data, and bulkheads so it cannot consume all my threads. The feature degrades and the core flow keeps working.

Common mistakes

  • No timeouts on remote calls.
  • Retrying without backoff, jitter or a limit.
  • Retrying at every layer of the stack.
  • A per-server rate limiter presented as a global one.
  • Sizing failover capacity so the survivors cannot take the load.
  • Quoting an availability figure without its downtime in minutes.
Header Logo