Practice · interactive
Design Lab
Pick a scenario, drag components onto the canvas and connect them. Then check your design: you'll get a rank, what you got right, what's missing and why, and the design an interviewer would hope to see.
See it in action: sizing
A photo-sharing app is about to hit its daily peak. Watch what the right sizing, and the wrong one, does to it.
Traffic–
p99 latency–
Errors–
Cost / month–
Press a button above: one run where it goes well, one where it goes wrong.
The brief
- A photo-sharing site growing from 1k to 1M daily users.
- Reads outnumber writes 50 : 1; photos are a few MB each.
- It must keep working when any single server dies.
The brief
- Turn long URLs into short codes and redirect anyone who opens one.
- 100M new URLs a month, read 100× more often than written.
- Redirects must be fast (under ~50 ms) and links must never break.
Full walkthrough: URL shortener (try the lab first!)
The brief
- Limit each API client to, say, 100 requests per minute, across dozens of gateway servers.
- Adds at most a millisecond or two to each request.
- Over-limit requests get 429 Too Many Requests.
Full walkthrough: Distributed rate limiter (try the lab first!)
The brief
- One-to-one and group chat with real-time delivery, like WhatsApp.
- 50M daily users, ~2B messages a day; history kept forever.
- Offline users get a push notification.
Full walkthrough: Chat system (try the lab first!)
The brief
- Users post; followers see a home timeline of recent posts.
- 200M daily users; timelines are read far more than posts are written.
- Timeline loads must be fast; a post can take a few seconds to appear.
Full walkthrough: News feed (like X / Twitter) (try the lab first!)
The brief
- Creators upload videos; viewers stream them smoothly on any connection.
- 500k uploads a day; at peak, millions of concurrent viewers.
- Videos can take a few minutes to become watchable.
Full walkthrough: Video streaming (like YouTube) (try the lab first!)
The brief
- Riders request rides; the system matches them with nearby available drivers.
- 1M active drivers send their location every 4 seconds.
- A driver must never be assigned to two rides at once.
Full walkthrough: Ride sharing (like Uber) (try the lab first!)
The brief
- Crawl 1 billion pages a month for a search index.
- Be polite: obey robots.txt, and never hammer one site.
- Don’t fetch the same URL twice.
Full walkthrough: Web crawler (try the lab first!)
The brief
- Any service can send push, SMS or email to users.
- 10M notifications a day, with bursts of millions during campaigns.
- Respect user preferences; one-time codes must never be delayed by marketing.
Full walkthrough: Notification system (try the lab first!)
The brief
- Suggest the top completions as someone types, like a search box.
- 5B suggestion requests a day; each answer in under 100 ms.
- Suggestions can be a day old (plus trending terms).
Full walkthrough: Search autocomplete (try the lab first!)
The brief
- Keep files in sync across a user’s devices; share folders.
- 50M users, ~10 GB each; small edits to big files shouldn’t re-upload everything.
- Never lose a file; never show a half-synced one.
Full walkthrough: File sync (like Dropbox) (try the lab first!)
Components
Drag onto the canvas, or tap to add.
Drag components here.
Then pull from a box's ● handle to another box to connect them.
Capacity calculator
How big does this design need to be? Set the traffic and the work each request does; the calculator sizes the servers, load balancers, database, cache and workers. It reads your canvas too: the database and load balancer you picked, and whether there's a cache.
Your estimate first
How many of each will this design need at peak? Guess, then reveal. Way too many and it blasts (wasted money); way too few and it melts down (overloaded).
The best design: Scale a web app to 1M users
- Users get static files and photos from the CDN and API calls through the load balancer.
- Stateless app servers behind the balancer scale out and survive single failures.
- Cache-aside in front of the database absorbs the 50 : 1 read load; read replicas add more.
- Photos live in object storage; the database only stores their keys.
The best design: URL shortener
- A load balancer fronts separate shorten and redirect paths.
- Redirects are a cache-aside lookup; most hit Redis and never touch the store.
- A key-value store (replicated) holds billions of code → URL rows.
- An ID generator hands out unique codes; clicks go to a queue for analytics.
The best design: Distributed rate limiter
- The limiter lives in the API gateway, so bad traffic is stopped before it reaches services.
- Counters sit in Redis, updated atomically with a Lua script so gateways never race.
- Rules come from a config service and are cached in each gateway.
The best design: Chat system
- Realtime gateways hold a connection per device; a session registry knows who is where.
- The chat service persists each message (wide-column store, partitioned by conversation) before acking.
- Offline users get push notifications.
The best design: News feed (like X / Twitter)
- Posting puts an event on a queue; fan-out workers push the post ID into followers’ timeline caches.
- Reading a timeline is one cache read, then a batch lookup to hydrate posts.
- Celebrities are the exception: their posts are merged in at read time (hybrid fan-out).
The best design: Video streaming (like YouTube)
- Uploads go straight to object storage; a queue triggers parallel transcoding workers.
- Viewers stream segments from the CDN, which pulls from storage on a miss.
- Metadata (status, renditions) lives in a database; the bytes never touch app servers.
The best design: Ride sharing (like Uber)
- Driver locations go to an in-memory geo index sharded by city: fast writes, nearest-driver queries.
- Matching reads the geo index and offers the ride to one driver at a time.
- Ride state lives in SQL with conditional updates, so a driver can’t be double-booked.
The best design: Web crawler
- The URL frontier (queue) balances priority and per-host politeness.
- Fetchers obey robots.txt, store pages in object storage, and hand pages to parsers.
- A Bloom filter drops URLs already seen; only new ones return to the frontier.
The best design: Notification system
- One notification API checks preferences, then enqueues.
- Separate queues per channel (and per priority) isolate slow providers and protect one-time codes.
- Workers retry with backoff and record results; idempotency keys prevent duplicates.
The best design: Search autocomplete
- The CDN answers popular prefixes; misses go to the suggest service.
- Suggestions come from an in-memory trie with top-k precomputed at every node.
- An offline aggregation job builds new trie snapshots from search logs.
The best design: File sync (like Dropbox)
- Clients split files into content-hashed chunks and upload only missing ones, directly to storage.
- The metadata service commits each version in a SQL transaction and appends to a change journal.
- A notification service tells other devices to fetch changes since their cursor.