Mid to Senior Engineer

System Design Interview Prep

A structured path from the interview framework through core concepts, key technologies and patterns to eighteen full problem breakdowns, each with diagrams and weak, solid and excellent answers to every deep dive.

Chapter 31 of 36Problem breakdowns · Design a Video Streaming Platform

Design a Video Streaming Platform

Video platforms are among the largest systems on the internet by data volume. The interview version asks you to design the pipeline that takes an uploaded video and delivers it smoothly to millions of viewers on every kind of device and network. The ideas that matter are the processing pipeline, adaptive bitrate streaming, and the economics of delivery through a content delivery network. This chapter focuses on those, building on the cloud drive chapter's treatment of large objects.

The chapter follows the usual shape: understand the problem, set up the interface, build the high-level design, then go deep on the questions interviewers use to separate levels.

1. Understanding the problem

Creators upload videos. Viewers search, browse and watch them. The platform must process each upload into a form that plays well everywhere, deliver it with little buffering, and keep costs under control.

Functional requirements

Core:

  1. Creators upload videos, including large ones.
  2. Viewers can watch videos on many devices with smooth playback.
  3. Viewers can search and browse, and resume where they left off.

Confirm in or out: live streaming, recommendations, comments, subtitles, monetisation, downloads for offline viewing and copyright detection. A sensible opening: "I will design upload, processing, playback and resume, and treat live streaming and recommendations as extensions."

Non-functional requirements

  • Smooth playback. Fast start and little buffering, on fast and slow networks.
  • Scale in bytes. Storage and delivery volumes are enormous.
  • Availability for playback, with eventual consistency for metadata such as view counts.
  • Cost efficiency. Storage and bandwidth dominate the bill, so every design choice has a cost.
  • Durability. An uploaded original must not be lost.

Estimation

Assume 200 million daily viewers each watching 40 minutes, and 500,000 hours of video uploaded per day.

QuantityCalculationResult
Watch time min minutes per day
Average delivery bitratesay 3 Mbit/sabout 22 MB per minute
Daily egress MBabout 180 PB per day
Egress rate180 PB / 86,400 sabout 2 TB per second
Upload500,000 hours per dayabout 6 hours uploaded per second
Raw storage per hoursay 3 GB at the source qualityabout 1.5 PB of originals per day

What the numbers say. Delivery volume is staggering: terabytes per second cannot come from a few data centres. It must be served from caches close to viewers. Storage grows at petabytes per day, so there are real decisions about which renditions to keep. And processing 500,000 hours of video a day is a massive, parallel compute workload.

Treat these as order-of-magnitude assumptions to state out loud and adjust, not as facts about any particular service.

2. The set up

Core entities

  • Video: identifier, owner, title, state (uploading, processing, ready, blocked), duration.
  • Rendition: one encoded version of the video at a resolution and bitrate.
  • Segment: a few seconds of one rendition.
  • Manifest: the file that describes the renditions and segments.
  • Watch progress: user, video, position.

API

POST /v1/videos/uploads              -> { upload_id, part_urls }   (resumable, chunked upload)
POST /v1/videos/{id}/complete        -> triggers processing
GET  /v1/videos/{id}/playback        -> { manifest_url, signed token }
PUT  /v1/progress/{video_id}         { position }                   (every few seconds)
GET  /v1/search?q=...                -> results

3. High-level design

Two main flows, with very different characteristics.

Upload and processing (write path). Heavy compute, asynchronous, tolerant of delay. The creator uploads the original to object storage in chunks. A pipeline splits it into segments, transcodes them in parallel into a ladder of resolutions and bitrates, generates audio tracks and thumbnails, runs checks, and publishes a manifest.

<!--fig:pipeline-->
upload Creator Original object storage Splitter segments Task queue Transcode 1080p Transcode 720p Transcode 480p Audio + thumbs Publish manifest + CDN Figure 1. Upload and processing: split into segments, transcode in parallel to a ladder of qualities, publish with a manifest.

Playback (read path). Latency-sensitive and enormous in volume. The player asks a playback API for a manifest, then fetches short segments from a CDN edge, adapting quality as conditions change.

<!--fig:playback-->
manifest 2 segments miss 3 events Player adaptive bitrate logic Playback API auth, manifest URL CDN edge caches segments Origin object storage Playback events quality, buffering Analytics + resume Figure 2. Playback: the player reads a manifest, fetches short segments from the nearest edge, and switches quality as bandwidth changes.

4. Potential deep dives

Deep dive 1: How do you process video at this scale?

The challenge. One long video, encoded at several qualities, is a lot of compute, and creators expect it to be playable soon.

Weak: transcode each video start to finish on one machine. A two-hour video at high quality takes a long time on one machine, and a failure restarts everything.

Solid: split into segments and transcode in parallel. Cut the video at keyframes into segments of a few seconds to a minute, put each (segment, quality) pair on a task queue, and let many workers encode them. A final step stitches the results and writes the manifest. A failed task is retried alone.

Excellent: parallelism, priority and early availability. Fan out across a worker fleet, and make tasks idempotent so retries and duplicates are safe. Use priority: encode the lowest-resolution rendition first, so the video becomes watchable quickly, then add higher qualities as they finish. Use per-title encoding: analyse each video's complexity and choose the bitrate ladder to fit, so a cartoon needs less bitrate than a sports match for the same quality. Scale workers on queue depth, using cheaper interruptible compute for work that tolerates retries. Track the state of each video as a workflow, with progress visible to the creator. Keep the original forever (or in cold storage), so you can re-encode when better codecs appear.

Deep dive 2: How does playback adapt to the network?

The challenge. Viewers have different devices and connections that change during playback.

Weak: serve one file at one quality. Slow connections buffer constantly, and fast ones get worse quality than they could.

Solid: adaptive bitrate streaming. Provide several renditions, each cut into short segments, described by a manifest. The player measures its download speed and buffer level and picks the best rendition it can sustain for the next segment, switching up or down between segments. Common protocols, HLS and MPEG-DASH, run over ordinary HTTP, so standard web caches can serve them.

Excellent: tuned for start time, stability and quality. Start at a moderate quality for a fast first frame, then climb. Keep a buffer target so short network dips do not stall playback, and avoid oscillating between qualities, which is more annoying than a steady lower quality. Prefetch the next segments. Use shorter segments for faster adaptation, balanced against request overhead. Choose codecs that balance compression and device support, offering a more efficient codec to devices that can decode it and a universally supported one as a fallback. Collect playback metrics (start time, rebuffer ratio, bitrate) from clients, and use them to tune the algorithm and detect regional problems.

Deep dive 3: How do you deliver terabytes per second?

The challenge. Origin storage cannot serve millions of concurrent viewers, and cross-continent latency would ruin playback.

Weak: serve everything from the origin. Bandwidth and latency make this impossible at scale.

Solid: a CDN. Cache segments at edge locations near viewers. Popular videos are served almost entirely from the edge, and only misses reach the origin.

Excellent: a tiered cache hierarchy and content placement. Use multiple tiers: edge caches near users, regional caches behind them, then the origin, so a miss in one edge is often a hit in the regional tier and the origin sees a tiny fraction of traffic. Viewing is extremely skewed: a small share of titles accounts for most hours. Pre-position the popular and the newly released titles on edge caches in advance, during off-peak hours, instead of waiting for first viewers to miss. For the long tail, accept an origin fetch on first request. Some large operators place cache servers inside internet providers' networks to cut transit costs and latency, and you can mention this as a strategy without claiming details of any company. Handle a cache stampede on a newly viral video with request coalescing at each tier. Protect content with signed URLs or tokens with short expiry, so only authorised viewers fetch segments. Monitor hit ratio by region, because a falling hit ratio is a cost and latency alarm.

Deep dive 4: Storage cost

The challenge. Every video is stored at many renditions, and the catalogue grows without end.

Weak: keep every rendition of every video forever on fast storage. The bill grows without bound, and most of it is for content nobody watches.

Solid: tier by popularity. Keep hot videos on fast storage and move cold ones to cheaper tiers.

Excellent: tier, prune and re-encode. Track watch frequency and move cold renditions to cheaper storage, delete the highest-bitrate renditions of rarely watched videos and regenerate them on demand from the original if the video becomes popular again. Keep the original in the cheapest durable tier. Choose codecs that reduce size for the same quality, accepting the one-time encoding cost, when a video will be watched many times. Compute the cost per hour stored and per hour delivered, and show where the money goes.

Deep dive 5: Resume position, view counts and metadata

Resume. Clients send the position every few seconds. Do not write each update to a durable database. Buffer in a fast store keyed by user and video, flush periodically, and tolerate losing the last few seconds if the store fails.

View counts. A popular video receives enormous numbers of view events. Do not increment a single row. Publish view events to a log, aggregate in streaming windows, and update the count periodically, showing slightly stale totals. Deduplicate and filter bots, and decide what counts as a view, such as watching for at least some seconds.

Search and browse. A search index fed from the metadata database with change capture, with ranking that blends text relevance with popularity and recency, as in the search engine chapter. Recommendations are a separate system: candidate generation, ranking and re-ranking, fed by watch history, and an extension unless asked.

Deep dive 6: Upload reliability and safety

Uploads. Use resumable, chunked uploads straight to object storage with signed URLs, as in the cloud drive chapter, so large files survive interruptions and do not pass through application servers. Verify integrity with checksums, and expire abandoned uploads.

Safety and rights. Content moderation and copyright checks run in the pipeline before publishing, using automated detection plus human review for flagged items. Treat this as a policy and legal area: say that rules vary by jurisdiction and that you would involve the relevant specialists. Make a takedown immediate by revoking playback tokens and purging caches.

Deep dive 7: Live streaming (an extension)

Live differs because there is no time to process offline. The broadcaster sends a continuous stream to an ingest server, which transcodes in real time into the ladder, cuts segments of a couple of seconds, and publishes them to the CDN as they are produced, with the manifest updated continuously. The trade-off is latency versus stability: shorter segments reduce delay and increase request overhead and risk of stalls. The edge fetches each new segment once and serves it to many viewers, so a million viewers do not mean a million origin requests. Plan for ingest redundancy, since losing the encoder loses the broadcast.

5. What is expected at each level

Mid-level. You describe uploading to object storage, processing into several qualities, a CDN for delivery, and a metadata database. You mention adaptive bitrate streaming.

Senior. You size the system, design parallel transcoding with segments and a queue, explain adaptive bitrate in detail, design CDN tiering and placement, and handle resume, counts and failures.

Staff. You reason about cost per hour, codec and ladder strategy, cache hit ratios and placement, quality-of-experience metrics, safety and rights, live latency trade-offs and how you would experiment with algorithm changes safely.

6. Interview questions and model answers

Q: Why split video into segments? It lets transcoding run in parallel across many workers with independent retries, and it enables adaptive streaming, where the player can switch quality between segments. Short segments are also cacheable as ordinary files at the CDN.

Q: How does adaptive bitrate streaming work? The video is encoded at several qualities and cut into short segments listed in a manifest. The player measures bandwidth and buffer, chooses a quality for each next segment and switches as conditions change.

Q: How do you serve millions of viewers at once? Through a tiered CDN. Popular segments are cached at the edge and in regional tiers, new releases are pre-positioned, and the origin sees only a small fraction of requests.

Q: How do you cut storage cost? Tier storage by popularity, delete high-bitrate renditions of cold videos and regenerate on demand, pick efficient codecs for popular content and keep originals in cheap durable storage.

Q: How do you count views at scale? Publish view events to a log and aggregate them in windows, with deduplication and bot filtering, and update the visible count periodically.

Q: What is hardest about live streaming? There is no offline processing, so latency and stability trade off through segment length, and the ingest and encoder are single points of failure that need redundancy.

7. Common mistakes

  • Serving video from the origin instead of a CDN hierarchy.
  • Transcoding a whole video on one machine.
  • Writing watch progress or view counts to a single database row per event.
  • Streaming one fixed quality.
  • Storing every rendition of every video at the highest tier forever.
  • No signed URLs, so content can be fetched without authorisation.
  • Forgetting that uploads should bypass application servers.
Header Logo