System Design Guide

Fundamentals · 14

Message queues and event streaming

Decouple services with queues and logs. Covers point-to-point vs pub/sub, Kafka-style logs, delivery guarantees, ordering, retries and dead-letter queues.

3 min read · 6 flashcards

A message queue sits between services: a producer writes a message and moves on, and a consumer processes it later. Work that doesn’t need to finish before the user gets a response, like sending emails, resizing images or updating feeds, belongs behind a queue.

API serverAPI serverQueue / topicWorker 1Worker 2Worker 3Dead-letterqueueafter N failures
Producers hand work to a queue; a consumer group shares it; failures go to a dead-letter queue

Why use a queue

  • Decoupling: producers don’t need to know who consumes, or whether they’re up right now.
  • Load levelling: a spike of 10,000 jobs waits in the queue instead of overwhelming workers.
  • Faster responses: the API returns as soon as the job is queued.
  • Scaling: add workers to drain faster.
  • Resilience: if a consumer crashes, messages wait instead of being lost.

Three styles

Style Delivery Retention Examples
Work queue (point-to-point) Each message to one consumer Deleted when processed SQS, RabbitMQ
Pub/sub Each message to every subscriber Usually short SNS, Google Pub/Sub, Redis Pub/Sub
Log / stream Consumers track their own position; many groups read the same data Kept for days or forever; replayable Kafka, Kinesis, Pulsar

A log like Kafka is append-only and split into partitions. Each consumer group reads every partition, and within a group each partition is read by one consumer. Because data is retained, you can add a new consumer later and replay history, which makes logs the backbone of event-driven systems and stream processing.

Delivery guarantees

  • At-most-once: may lose messages, never duplicates (acknowledge before processing).
  • At-least-once: never loses, may duplicate (acknowledge after processing). This is the common default.
  • Exactly-once: very hard end to end. In practice you get effectively once through at-least-once delivery plus idempotent consumers: dedupe by message ID, or design updates so applying them twice is harmless (SET status = 'paid', not balance += 10).

Ordering

  • Global ordering doesn’t scale. Kafka orders messages within a partition.
  • Choose a partition key that groups messages needing order: all events for order_id = 42 go to the same partition, so they’re processed in order.

When things fail

  • Retries with backoff for transient errors (see reliability).
  • Visibility timeout (SQS): a message is hidden while being processed and reappears if the worker dies.
  • Dead-letter queue (DLQ): after N failed attempts, park the message for inspection instead of blocking the queue forever.
  • Backpressure: watch queue depth and consumer lag; scale consumers or slow producers when they grow.

Pros

  • Decouples services; failures don’t cascade immediately
  • Absorbs traffic spikes
  • Enables async processing and fan-out
  • Logs give replay, audit and stream processing

Cons

  • Eventual: results aren’t immediate
  • Duplicates and reordering to handle
  • More infrastructure to operate and monitor
  • Harder debugging: requests become chains of events

Test yourself

Answer in your head, then click a card to check. All cards are in the Anki deck.