System Design Guide

Fundamentals · 04

Scaling: vertical vs horizontal

Two ways to handle more load (a bigger machine or more machines), why horizontal scaling needs stateless servers, and where state should live instead.

3 min read · 5 flashcards

Every popular system eventually outgrows one machine. There are two ways to respond, and real systems use both.

Vertical scaling (scale up)

Give the machine more CPU, RAM or faster disks.

Pros

  • No code changes: the simplest option by far
  • No distributed-systems problems (no network splits, no consistency issues)
  • Great for databases early on: one big box goes a long way

Cons

  • Hard ceiling: the largest machine available
  • Cost grows faster than capacity at the high end
  • Still a single point of failure
  • Upgrades often mean downtime

Horizontal scaling (scale out)

Add more machines and spread the load across them with a load balancer.

ClientsLoad balancerApp server 1App server 2App server 3Session store(Redis)Database
Horizontal scaling: identical stateless servers, state kept elsewhere

Pros

  • Practically unlimited capacity: add commodity machines
  • Fault tolerant: lose one server and the rest carry on
  • Scale up and down with demand (autoscaling)

Cons

  • Needs a load balancer and stateless servers
  • Distributed-systems complexity: partial failures, consistency
  • Data tiers are much harder to scale out than app tiers

Stateless services: the key to scaling out

A server is stateless if any instance can handle any request, so nothing about a user lives only in one server’s memory. If server 2 holds Alice’s session and dies, Alice is logged out. Worse, the load balancer must keep sending her to server 2 (“sticky sessions”), which spreads load unevenly.

Where state goes instead:

  • Shared session store: Redis or Memcached, keyed by session ID.
  • Client-side tokens: a signed JWT carries the user’s identity, so the server checks the signature and needs no lookup. The cost is that revoking a token before it expires needs extra work, like a deny-list or short lifetimes.
  • Databases and object storage: for anything durable, like uploads or carts.

Once servers are stateless you can autoscale: add instances when CPU or request rate rises, and remove them when it falls.

Scaling the data tier

App servers are easy; data is hard. In rough order:

  1. Vertical scaling of the database. Often enough for a long time.
  2. Caching to absorb repeated reads.
  3. Read replicas to spread reads.
  4. Sharding to spread writes and data across machines. Powerful, but complex.

Test yourself

Answer in your head, then click a card to check. All cards are in the Anki deck.