Mid to Senior Engineer

System Design Interview Prep

A structured path from the interview framework through core concepts, key technologies and patterns to eighteen full problem breakdowns, each with diagrams and weak, solid and excellent answers to every deep dive.

Chapter 7 of 36Key technologies · Redis

Key Technology: Redis

Redis is an in-memory data store that shows up in a large fraction of system design answers: as a cache, a rate limiter, a leaderboard, a session store, a lock service. Interviewers do not expect you to know every command. They expect you to know what Redis is good at, how it scales and fails, and where it is the wrong choice. This chapter gives you that, so that when you write "Redis" on a whiteboard you can defend it.

Version and product details change over time, so where a specific number or feature matters, treat it as something to confirm against the current documentation of the version you would use.

1. What Redis is

Redis keeps its data in memory, which is why reads and writes take well under a millisecond on the server. It is more than a key-value cache: values are data structures that the server understands and can modify atomically.

StructureWhat it gives youTypical use
StringA value up to a large size, or an integer you can incrementCached objects, counters
HashA small map of fields under one keySession or user objects
ListAn ordered list with push and pop at both endsSimple queues, recent items
SetUnique members, with union and intersectionTags, unique visitors
Sorted setUnique members ordered by a scoreLeaderboards, time-ordered indexes
StreamAn append-only log with consumer groupsEvent pipelines
Bitmap, HyperLogLogCompact flags, approximate distinct countsDaily active users
GeospatialMembers indexed by coordinatesNearby search

Every key can have a time to live, after which it is removed automatically. That single feature is behind most uses: sessions that expire, rate-limit windows that reset, cache entries that go stale.

<!--fig:uses-->
Redisin-memory data structures Cachestrings, hashes + TTL Rate limiteratomic counters, scripts Leaderboardsorted sets Session storehashes + expiry Distributed lockSET key NX with expiry Queues + pub/sublists, streams, channels Figure 1. What Redis is used for in designs: each use maps to a data structure it already provides.

2. How it works, and why it is fast

Memory first. Data lives in RAM, so there is no disk read on the hot path.

A simple execution model. Command execution is, in the core design, handled by a single thread, so each command runs to completion without locks. That makes individual commands atomic, and it means one slow command, such as scanning a huge key, blocks everyone behind it. Newer versions use extra threads for network input and output, but the principle that commands are executed one at a time is what you should remember.

Persistence is optional and secondary. Two mechanisms:

  • A snapshot writes the whole dataset to disk periodically. Recovery is fast, and you lose writes since the last snapshot.
  • An append-only log records every write command. You can choose how often it is flushed to disk, trading speed for safety, and it can be rewritten to stay compact.

Many deployments use Redis purely as a cache with persistence off, and accept that a restart starts empty. Others use both mechanisms because the data is hard to rebuild. Say which you choose and why.

Throughput. A single node commonly handles on the order of a hundred thousand simple operations per second. Treat that as a rough planning number to verify with a benchmark on your hardware and value sizes, not as a guarantee.

3. Scaling and availability

Replication and failover

A primary accepts writes and streams them to replicas. Replicas can serve reads, and one of them takes over if the primary fails. Replication is asynchronous by default, so a failover can lose the writes the replica had not yet received.

A failover coordinator (called Sentinel in the standard distribution) watches the primary, and a majority of its members must agree that it is down before a replica is promoted. The majority rule prevents a partitioned minority from promoting its own primary.

<!--fig:ha-->
commands async stream persist health promote Application client library Primary reads + writes Replica 1 Replica 2 Failover coordinator majority vote Disk snapshot + log Figure 2. Primary with replicas; a failover coordinator promotes a replica. Replication is asynchronous, so a failover can lose the latest writes.

Partitioning with Redis Cluster

One node is limited by memory and by single-threaded throughput. Redis Cluster splits the key space into 16,384 hash slots. A key's slot is a hash of the key modulo 16,384, and each primary owns a range of slots. A client library caches the slot map and sends each command to the owning node. If the map is stale, the node replies with a redirect and the client updates its map. Moving a slot between nodes rebalances the cluster.

Two consequences matter in interviews:

  • Multi-key operations only work on keys in the same slot. A transaction or script touching keys on different nodes is rejected. You can force related keys into one slot with a hash tag, a part of the key name in braces that is hashed instead of the whole key.
  • A hot key lives on one node. Partitioning spreads keys, not the traffic to a single key.

4. Interview use cases, with the reasoning

Use case 1: Cache

Store computed or fetched objects under a key with a TTL. Use the cache-aside pattern from the scaling chapter: read the cache, fall back to the database on a miss, then fill. Configure an eviction policy for when memory is full. The common choices evict the least recently used or least frequently used keys, either among all keys or only among keys that have a TTL. Choose "all keys" for a pure cache.

What to say: "Redis as a cache, cache-aside, a TTL on every entry, an LRU-type eviction policy, and the database is the source of truth, so losing the cache is slow but not harmful."

Use case 2: Rate limiting

Counters with expiry are a natural fit. A fixed-window limiter increments a key that includes the window, such as the user and the current minute, and sets it to expire. A token bucket needs a read-modify-write, which must be atomic.

Weak: read the counter in the application, compare, then write it back. Two servers can both read "one token left" and both spend it. Solid: use the atomic increment command, which returns the new value, and compare it to the limit. Excellent: put the whole algorithm in a server-side script, which Redis runs without interruption. The script reads the bucket, refills it from the elapsed time, spends a token and writes it back in one atomic step, and returns the decision. This is the standard answer for a token bucket on a shared store.

Use case 3: Leaderboard

A sorted set keeps members ordered by score. Adding or updating a score costs , and fetching the top ten or a player's rank is also logarithmic. A real-time game leaderboard for millions of players is a few commands: set the score, read the top range, read a member's rank.

What to watch: a single sorted set lives on one node. For a very large leaderboard, partition by season or region, or shard the set and merge the top results.

Use case 4: Session store

Store the session as a hash under a key containing the session identifier, with a TTL that is refreshed on activity. Any app server can read it, so the application tier stays stateless.

What to watch: if sessions are lost on a failover, users are logged out. Decide whether that is acceptable. If not, enable persistence and replicas, or keep the session in a signed token.

Use case 5: Distributed lock

A common pattern sets a key only if it does not exist, with an expiry, using one atomic command, and releases it by deleting the key only if the stored value is the owner's unique token.

Weak: a lock with no expiry. If the holder crashes, the lock is never released. Solid: a lock with an expiry and a unique owner token, released only by the owner. Excellent: acknowledge the limits. If the holder pauses longer than the expiry, for example during a garbage collection pause, another process takes the lock and the first one resumes believing it still holds it. A lock with an expiry therefore cannot guarantee mutual exclusion on its own. If correctness depends on it, use a fencing token, an increasing number checked by the protected resource, or a consensus-based coordination service. Also, a lock held on a single primary can be lost in a failover before it is replicated. Algorithms exist that acquire a lock across several independent nodes, and there is a well-known debate about whether they give strong enough guarantees. The honest interview answer is that Redis locks suit efficiency (avoid doing duplicate work) better than strict correctness (never allow two writers).

Use case 6: Queues, streams and publish-subscribe

Lists can act as a simple work queue. Streams provide an append-only log with consumer groups, acknowledgements and the ability to claim messages from a failed consumer, which is close to a lightweight Kafka. Publish-subscribe channels deliver a message to current subscribers only, with no persistence: a subscriber that is offline misses the message.

What to say: use streams for modest-scale pipelines inside a system that already has Redis, and use a dedicated log such as Kafka for large, durable, replayable event streams.

5. Where Redis is the wrong choice

  • Large datasets that do not fit in memory economically. RAM costs far more than disk, so a multi-terabyte, rarely accessed dataset belongs in a disk-based store.
  • The only copy of important data. Replication is asynchronous, and even with persistence a failure can lose recent writes. Keep a durable source of truth.
  • Complex queries. There are no joins or ad hoc queries. Secondary lookups need extra structures you maintain yourself.
  • Strong consistency across keys. Cluster mode limits multi-key operations to one slot, and replication can lose writes.
  • Very large values or heavy scans. One slow command blocks the single execution thread.

6. Failure modes and operations

  • Memory exhaustion. When memory is full, Redis either evicts keys or rejects writes, depending on the policy. Monitor memory use and evictions.
  • Hot keys. Detect them from access sampling, and spread reads with replicas, a local in-process cache, or key splitting.
  • Cache stampede. Many requests miss together when a popular key expires. Use request coalescing, randomised TTLs, and refresh before expiry.
  • Big keys. A set or list with millions of members makes some commands slow and a deletion a long pause. Split such keys and use incremental scans, not commands that return everything.
  • Persistence overhead. Snapshots fork the process, which can double memory use under heavy writes. Leave headroom.
  • Failover data loss. Mention that asynchronous replication can lose acknowledged writes, and that a "wait for replicas" option reduces but does not eliminate the risk.

7. Interview questions and model answers

Q: Why is Redis fast? Data is in memory, commands are simple and run one at a time without locks, and the protocol is lightweight. A single node commonly serves on the order of a hundred thousand operations per second.

Q: How would you use Redis to rate limit an API? A shared counter or bucket per key, updated atomically. For a token bucket I would run the refill-and-spend logic in a server-side script so that concurrent gateways cannot both spend the last token. I would shard keys across a cluster and decide whether to fail open if Redis is unreachable.

Q: How does Redis scale beyond one machine? Redis Cluster partitions keys over 16,384 hash slots across primaries, each with replicas. Clients route by slot. Multi-key operations need all keys in one slot.

Q: Is a Redis lock safe? It is good for avoiding duplicate work, but a lock with an expiry can be held by two clients after a pause or a failover. For correctness I would add a fencing token or use a consensus-based service.

Q: When would you not use Redis? For the only copy of important data, for datasets much larger than affordable memory, or when I need rich queries or strong consistency across keys.

Q: What happens to the data when a primary fails? A replica is promoted after a majority agrees. Because replication is asynchronous, the most recent writes can be lost. For a cache that is fine, because the data can be reloaded from the database.

8. Common mistakes

  • Using Redis as the only store for data that must not be lost.
  • A non-atomic read-then-write for counters and limits.
  • A distributed lock with no expiry, or with no discussion of pauses.
  • No eviction policy or memory headroom.
  • Ignoring hot keys and big keys.
  • Using publish-subscribe where messages must not be lost.
  • Multi-key operations in cluster mode without considering slots.
Header Logo