Mid to Senior Engineer

System Design Interview Prep

A structured path from the interview framework through core concepts, key technologies and patterns to eighteen full problem breakdowns, each with diagrams and weak, solid and excellent answers to every deep dive.

Chapter 22 of 36Problem breakdowns · Design a Chat System

Design a Chat System

A chat system is where the stateless-servers rule breaks. To push a message the instant it is sent, a server must hold an open connection to the recipient, which makes that server stateful. How you route to the right connection, store history, order messages and guarantee delivery is the heart of the design, and it is a favourite question because every part of it has a trap.

This chapter builds the design in the usual order, then goes deep on delivery guarantees, ordering, group chat and presence. Each deep dive compares a weak, a solid and an excellent answer.

1. Understanding the problem

A messaging product lets people exchange text, and often media, in one-to-one and group conversations, in real time.

Functional requirements

Core:

  1. Users send and receive one-to-one messages in real time.
  2. Users can create group conversations and message them.
  3. Messages are stored, so a user who was offline receives them later and can scroll through history.

Confirm in or out: delivery and read receipts, media attachments, multiple devices per user, online status, end-to-end encryption, search. A sensible opening: "I will build one-to-one and group messaging with durable storage and delivery status, then discuss presence, devices and encryption."

Ask about group size. A group of 200 and a broadcast channel with 10 million subscribers are different designs. Assume groups are capped at a few hundred members.

Non-functional requirements

  • Low latency. A message appears within a few hundred milliseconds when both users are online.
  • No lost messages, and ordering within a conversation. These are the properties users notice.
  • High availability, including graceful behaviour on flaky mobile networks.
  • Scale. Assume 500 million daily users.

Estimation

Assume each user sends 40 messages a day.

QuantityCalculationResult
Messages per day, about 230,000 per second
Storage100 bytes per messageabout 2 TB per day, 730 TB per year before replication
Concurrent connectionssay 100 million at peakthousands of connection servers

A well-tuned server can hold tens of thousands to a few hundred thousand idle connections depending on memory and the stack. Assume 50,000 per server, which gives about 2,000 connection servers. Present this as a figure you would measure, not a constant.

What the numbers say. The message rate is high but manageable with partitioned storage. The number of open connections, not the message rate, sizes the real-time tier, and that tier is stateful.

2. The set up

Core entities

  • User and device.
  • Conversation (one-to-one or group) with its list of members.
  • Message: conversation, sender, content, a sequence number within the conversation, timestamp.

API and protocol

Most operations are ordinary HTTP: log in, list conversations, load message history, upload media. The real-time channel is a WebSocket, a persistent two-way connection, over which the client sends messages and the server pushes them.

client -> server   { "type": "send", "conversation": "c9", "client_msg_id": "u1-77", "text": "hi" }
server -> client   { "type": "message", "conversation": "c9", "seq": 412, "from": "u2", "text": "hi" }
server -> client   { "type": "ack", "client_msg_id": "u1-77", "seq": 412 }

The client-generated message identifier is what makes retries safe: if the client resends after a timeout, the server recognises the identifier and does not store a duplicate.

How the client stays connected

TechniqueHow it worksVerdict for chat
PollingClient asks every few secondsWasteful, high latency
Long pollingRequest held until data or timeoutWorks everywhere, a reconnect per message
Server-sent eventsOne-way stream from serverServer to client only
WebSocketPersistent two-way connectionBest fit, needs stateful servers

Use WebSockets with a long-polling fallback. Clients send heartbeats so both sides detect dead connections, and reconnect with exponential backoff and jitter, because a server restart would otherwise make all of its clients reconnect at the same instant.

3. High-level design

The components:

  • Gateway (connection) servers terminate WebSockets. They are stateful but thin: authenticate, hold the connection, and forward messages in both directions.
  • A session registry, a fast shared store, maps each user to the gateway currently holding their connection.
  • The chat service receives each message, assigns it an order, persists it and routes it.
  • A message store, partitioned by conversation.
  • A push notification service for users who are offline.
<!--fig:hld-->
WebSocket WebSocket persist lookup if offline Alice Bob Gateway 1 WebSockets Gateway 2 WebSockets Chat service id, order, route Message store partitioned by conversation Session registry user to gateway Push notificationservice Figure 1. Gateways hold the WebSocket connections; the session registry says which gateway holds each user.

The life of a one-to-one message

  1. Alice's client sends the message over her WebSocket to her gateway.
  2. The gateway passes it to the chat service.
  3. The chat service assigns a sequence number within the conversation and writes the message durably.
  4. It acknowledges to Alice. Her client shows one tick: sent. The write happens before the acknowledgement, so an acknowledged message is never lost.
  5. The chat service asks the session registry where Bob is connected.
  6. If Bob is online, it forwards the message to his gateway, which pushes it down his socket. Bob's client acknowledges receipt, giving Alice two ticks: delivered.
  7. If Bob is offline, the chat service sends a push notification and keeps the message pending. When Bob reconnects, his client syncs every message after the last sequence number it holds.
<!--fig:flow-->
1 send 2 3 persist (seq n) 4 ack: sent 5 where is Bob? 6 push 7 Alice Gateway A Chat service Gateway B Bob Message store Session registry Bob offline: send a push notification, keep the messagestored, and let his client sync 'after seq n' on reconnect. Figure 2. Life of a message: persist first, acknowledge, then deliver. Reconnecting clients sync from their last sequence number.

Notice what carries the correctness. Real-time push is a best-effort optimisation. Correctness comes from durable storage plus the sync-on-reconnect step, so even a missed push is recovered.

4. Potential deep dives

Deep dive 1: How do you guarantee messages are never lost?

The challenge. Networks drop connections, servers crash, and phones go in and out of coverage. The user's expectation is that a message acknowledged as sent will arrive.

Weak: forward the message and hope. The chat service relays the message to the recipient's gateway and returns. If the gateway crashes, or the recipient's connection dropped a moment ago, the message vanishes and nobody knows.

Solid: persist before acknowledging, and ack at each hop. Write the message to the store first and only then acknowledge to the sender. Delivery to the recipient is acknowledged by the recipient's client, and unacknowledged messages are retried. The sender sees status: sent, delivered, read.

Excellent: durable log plus client sync. Treat the stored, ordered conversation as the source of truth and live push as an accelerator. Each client remembers the highest sequence number it has received per conversation. On reconnect, or on a periodic check, it asks "give me everything after sequence n" from durable storage. A lost push, a crashed gateway or a long offline period all heal the same way, with no special cases. Add client-generated identifiers so retries are idempotent, and the guarantee becomes at-least-once delivery with deduplication, which behaves like exactly-once for the user.

Deep dive 2: How do you order messages?

The challenge. Two people type at once, messages travel different paths, and device clocks disagree. Everyone in the conversation should see the same order.

Weak: order by timestamps from the sender's device. Device clocks are wrong, sometimes by minutes, and a user can set theirs on purpose. Messages appear out of order or in the future.

Solid: order by server arrival time. The server stamps messages on arrival. It is better, but with several chat service instances the stamps still come from different machines whose clocks differ slightly, and arrival order can differ from the order the sender intended.

Excellent: a sequence number per conversation, assigned by the partition that owns the conversation. Global ordering is neither cheap nor needed. You need order within a conversation. Route all writes for a conversation to one partition, which assigns a monotonically increasing sequence number. Clients display and sync by sequence number, never by timestamp. Because one partition owns the counter, there is no coordination across machines on the hot path, and "latest 50 messages" becomes a single-partition range read in order. State the cost: a very busy conversation is limited by the throughput of one partition, which is the hot-partition problem discussed under group chat.

Deep dive 3: How do you route a message to the right connection?

The challenge. Alice's gateway and Bob's gateway are different machines among thousands, and Bob's connection can move.

Weak: broadcast to every gateway. Every message is sent to all gateways, each checking whether it holds the recipient. It works for ten servers and collapses at two thousand.

Solid: a session registry. When a client connects, its gateway records user -> gateway in a fast shared store, with an expiry refreshed by heartbeats. The chat service looks the recipient up and forwards to exactly one gateway. If the entry is missing or stale, treat the user as offline and fall back to push plus sync.

Excellent: registry plus resilience. The same, with: several connections per user (a user has a phone and a laptop), so the registry holds a set; short entry lifetimes so a crashed gateway's entries expire on their own; and a retry path, because the registry can say "online" for a user whose connection has just dropped. Delivery failure is not an error. It simply means the stored message waits for sync. Optionally, partition the chat service by conversation so that all messages of one conversation are handled by one instance, which can cache the membership and registry lookups.

Deep dive 4: Group chat

The challenge. One message must reach many members, who may be offline or on several devices.

Weak: copy the message into every member's inbox. Storage grows with group size, and a message to a group of 10,000 causes 10,000 writes and 10,000 deliveries at once.

Solid: store the message once, deliver to each member. Keep one copy in the group's conversation partition with a sequence number. Each member tracks their own last-read position, so read state is a small per-member pointer, not a message copy. Deliver in real time to members who are online, and notify those who are not.

Excellent: scale delivery by audience size. For small groups, push to every online member as above. For very large groups or channels, pushing to everyone at once overloads the system, so treat them like the feed problem. Deliver to currently connected subscribers through the gateways that hold them, let everyone else pull on open using sequence numbers, and send aggregated notifications ("37 new messages") instead of one per message. Keep the member list in a fast store and apply joins and leaves from a given sequence number onward, so the history a new member can see is well defined. A hot conversation can be split into time buckets so no single partition holds it forever.

Deep dive 5: Presence and "last seen"

The challenge. Online status looks trivial and generates enormous update volume.

Weak: broadcast every status change to all of a user's contacts. A user with 500 contacts who toggles between foreground and background causes hundreds of messages each time, and nearly nobody is looking.

Solid: heartbeats with expiry. The client sends a heartbeat every few seconds, the presence service stores last_seen with a short expiry, and a user is online while the entry is fresh. Apply a small grace period before declaring a user offline, so brief network drops do not cause flapping.

Excellent: subscribe only where it matters. Send presence updates only to people currently viewing that contact, such as those with the conversation open, and fetch it on demand with a short cache otherwise. Coalesce rapid changes, and make presence best-effort: it is allowed to be slightly wrong, and it must never delay message delivery.

5. Multiple devices, security and failure

Multiple devices. The session registry holds several connections per user. A message is delivered to all of them, and each device syncs independently from its own last sequence number. Read state is synchronised across devices through the server.

Security. Use TLS for transport and authenticate the WebSocket at connection time with a short-lived token. With end-to-end encryption, the sender encrypts for each recipient device so servers carry only ciphertext. Mention the consequences: the server cannot search message content, and group messaging and multiple devices require key management for every device and every membership change.

Failure handling.

  • A gateway crashes: its clients reconnect to other gateways, re-register in the registry and sync from their last sequence number.
  • The registry says online but delivery fails: store and notify, and let sync recover.
  • A chat service node dies after storing but before acknowledging: the client retries, and the client-generated identifier prevents a duplicate.
  • A reconnect storm after a restart: exponential backoff with jitter on the client, and connection limits per gateway during recovery.

6. What is expected at each level

Mid-level. You get to WebSockets for push, a database for history, and a way to deliver offline messages. You can say why messages need a server-assigned order when prompted.

Senior. You separate the live path from the durable path, explain sync-on-reconnect, design the session registry and ordering per conversation, and handle at-least-once delivery with idempotent identifiers. You size the connection tier from your estimate.

Staff. You discuss large-scale behaviour: reconnect storms, hot conversations, multi-region placement of users, encryption trade-offs, and how presence and typing indicators are kept from overwhelming the system. You also say what you would not build: for example, global ordering.

7. Interview questions and model answers

Q: Why are chat servers stateful, and how do you scale them? Each holds long-lived WebSocket connections. I scale by adding gateways, keeping them thin, and tracking which user is on which gateway in a shared session registry. The durable state lives in the message store, so a gateway can die without data loss.

Q: How do you guarantee a message is not lost? Persist before acknowledging, and treat live delivery as best-effort on top of a durable store. A reconnecting client syncs everything after its last sequence number, so even a missed push is recovered.

Q: How do you order messages? Per conversation, with a monotonically increasing sequence number assigned by the partition that owns the conversation. I do not use sender clocks, and I do not attempt global ordering.

Q: How would you design group chat for 100,000 members? Store each message once, track read position per member, deliver in real time to those online, let the rest pull on open, and aggregate notifications. Per-recipient fan-out would overload the system.

Q: How do you show who is online without melting the system? Heartbeats with short-lived entries, updates sent only to users currently viewing the contact, and a grace period to damp flapping.

8. Common mistakes

  • Acknowledging a send before it is durably stored.
  • Treating the live push as the source of truth, with no sync path.
  • Relying on timestamps from clients for ordering.
  • Broadcasting presence to every contact.
  • Forgetting reconnect storms after a server restart.
  • Ignoring multi-device behaviour.
  • Copying each group message into every member's inbox.
Header Logo