System Design Guide

Free study guide · system design interviews

Learn to design large-scale systems from scratch.

Every building block explained plainly: what it is, when to use it and what it costs, with a diagram for each. Then ten classic interview questions worked through step by step, showing the thinking and the trade-offs, not just the final boxes and arrows.

19 fundamentals 10 worked questions 149 flashcards

A study plan

Roughly three to four weeks at an hour a day. Tick pages off with “Mark as done” at the bottom of each one; your progress is saved in this browser.

  1. Learn the method. Read how to approach the interview and back-of-the-envelope estimation. Every answer in this guide follows that framework.
  2. Foundations. Networking, scaling and load balancing: how a request reaches your servers, and how you add more of them.
  3. Data. Databases, indexing, replication, sharding, consistent hashing and the CAP theorem. This is where most interview trade-offs live.
  4. Speed and decoupling. Caching, CDNs, message queues, APIs and rate limiting.
  5. The rest of the toolkit. Object storage, unique IDs and reliability patterns.
  6. Practise out loud. For each question, set a 45-minute timer and sketch your own answer before reading the walkthrough. Then compare.
  7. Review daily. Ten minutes of flashcards (or the Anki deck) keeps the terms and numbers fresh.

Fundamentals

Each page covers what the concept is, how it works, the pros and cons, and what interviewers listen for.

01

The interview framework

A repeatable, seven-step framework for a 45-minute design interview, and what interviewers are actually scoring while you talk.

02

Estimation

Turn "100 million users" into requests per second, storage and bandwidth in a minute of arithmetic, and use the numbers to decide what's hard.

03

Networking basics

What happens between typing a URL and seeing a page, and the networking choices (TCP vs UDP, HTTP versions, proxies) that show up in every design.

04

Scaling basics

Two ways to handle more load (a bigger machine or more machines), why horizontal scaling needs stateless servers, and where state should live instead.

05

Load balancing

How load balancers spread traffic across servers, the difference between layer 4 and layer 7, the main algorithms, and how to keep the balancer itself from becoming a single point of failure.

06

Caching

Keep hot data in fast memory to cut latency and database load, and the strategies, eviction policies and failure modes (stampedes, stale data, hot keys) that come with it.

07

CDNs

Serve static and cacheable content from servers near your users, using pull or push models, cache headers and versioned URLs.

08

SQL vs NoSQL

The main database families, ACID versus BASE, and how to pick one from the access patterns rather than from fashion.

09

Indexes & storage engines

How databases find rows fast (B-trees and LSM-trees), the trade-off between read and write speed, and how to design indexes for your queries.

10

Replication

Keep copies of data on several machines for availability, read scaling and durability. Covers leader–follower, multi-leader and leaderless replication, and the lag problems each brings.

11

Sharding

Split a dataset across many machines so writes and storage scale out. Covers range, hash and directory sharding, choosing a shard key, hotspots and resharding.

12

Consistent hashing

Map keys to nodes so that adding or removing a node moves only a small fraction of keys, using a hash ring and virtual nodes.

13

CAP & consistency

What CAP really says, why it's about network partitions, how PACELC adds the everyday latency trade-off, and the consistency models between "strong" and "eventual".

14

Queues & streams

Decouple services with queues and logs. Covers point-to-point vs pub/sub, Kafka-style logs, delivery guarantees, ordering, retries and dead-letter queues.

15

APIs & real-time

REST, gRPC and GraphQL compared; pagination, idempotency and versioning; and how servers push updates with long polling, SSE, WebSockets and webhooks.

16

Rate limiting

Protect services from overload and abuse by capping requests per client. Covers the five classic algorithms, where to enforce limits, and doing it across many servers.

17

Blob storage

Where to put photos, videos and files, how clients upload large files directly with pre-signed URLs, and how object stores stay durable and cheap.

18

Unique IDs

Create IDs across many machines without collisions, from UUIDs to Twitter's Snowflake, and what each option means for ordering, size and coordination.

19

Reliability patterns

The patterns that keep a distributed system up when its parts fail. Covers redundancy, timeouts, retries with backoff and jitter, circuit breakers, bulkheads, idempotency, graceful degradation and observability.

Practice questions

Full walkthroughs: requirements, estimates, API, data model, high-level design, deep dives and failure modes, with the reasoning spoken out loud.

Q1 · Easy

URL shortener

Turn long URLs into short codes and redirect at high read volume. A classic warm-up that tests ID generation, caching and read-heavy design.

Q2 · Medium

Rate limiter

Limit requests per user, API key or IP across a fleet of gateway servers with minimal added latency. Tests algorithms, atomicity in a shared store and failure trade-offs.

Q3 · Hard

Chat system

Real-time one-to-one and group messaging with delivery receipts, offline delivery and multi-device sync. Tests WebSockets at scale, routing, ordering and delivery guarantees.

Q4 · Hard

News feed

Users post, follow others and read a home timeline. The classic fan-out problem, including the celebrity case that breaks the simple solution.

Q5 · Hard

Video streaming

Upload, transcode and stream video to millions of viewers with smooth playback. Tests media pipelines, adaptive bitrate streaming and CDN-heavy design.

Q6 · Hard

Ride sharing

Match riders to nearby drivers in seconds while tracking a million moving cars. Tests geospatial indexing, high-frequency writes and safe concurrent assignment.

Q7 · Medium

Web crawler

Download a billion pages a month politely and without repeats. Tests queue design (the URL frontier), deduplication with Bloom filters and partitioned, fault-tolerant workers.

Q8 · Medium

Notification system

Send push, SMS and email notifications for many services, reliably and without spamming users. Tests queues per channel, retries with idempotency, preferences and third-party providers.

Q9 · Medium

Search autocomplete

Suggest the top completions as someone types, in under 100 ms, for billions of queries a day. Tests tries with precomputed top-k, offline data pipelines and aggressive caching.

Q10 · Hard

File sync (Dropbox)

Store files in the cloud and keep them in sync across devices, efficiently and without losing edits. Tests chunking and deduplication, metadata consistency, sync protocols and conflicts.

Review with flashcards

Key terms and concepts from every page, as an in-browser quiz and a downloadable Anki deck for spaced repetition.