System Design Guide

Practice · Question 5 · Hard

Design a video streaming platform (like YouTube)

Upload, transcode and stream video to millions of viewers with smooth playback. Tests media pipelines, adaptive bitrate streaming and CDN-heavy design.

5 min read · 4 flashcards

1. Clarify requirements

Functional

  • Users upload videos (up to a few GB).
  • Users watch videos with smooth playback on any device and connection.
  • Basic metadata: title, description, thumbnails; view counts.
  • Out of scope: recommendations, comments, live streaming, monetisation.

Non-functional

  • Smooth playback: fast start, rare buffering.
  • Global audience, with huge read bandwidth.
  • Durable storage: never lose an uploaded video.
  • Uploads can take minutes to become watchable, which is acceptable.
  • Scale: 100M DAU, ~5 videos × 5 minutes watched per user per day; 500k uploads/day.

2. Estimates

  • Watch time: 100M × 25 min ≈ 2.5B minutes/day. At peak, say 5M concurrent viewers × ~5 Mbps ≈ 25 Tbps of egress. That’s only possible with a large CDN.
  • Uploads: 500k/day ≈ 6 uploads/s; × ~300 MB raw ≈ 150 TB/day of raw video.
  • Transcoded output: several resolutions (240p–4K) and codecs, roughly 1–2× the raw size, so expect PB-scale storage per year. That calls for object storage with lifecycle tiers.

3. API

POST /v1/videos                    { title, description, size }
   → { video_id, upload: { multipart pre-signed URLs } }
POST /v1/videos/{id}/complete      (all parts uploaded)
GET  /v1/videos/{id}               → { title, status, thumbnails, manifest_url }
GET  {cdn}/videos/{id}/master.m3u8 → adaptive streaming manifest
POST /v1/videos/{id}/views         (batched / sampled)

4. Data model

Data Store
videos(video_id, owner_id, title, description, status, duration, created_at) SQL (sharded by video_id) + cache
renditions(video_id, resolution, codec, bitrate, manifest_key) SQL
Raw uploads, segments, thumbnails Object storage, keyed videos/{id}/{rendition}/{segment}.ts
View counts Counters aggregated by a stream processor into a KV store

5. High-level design

CreatorUpload APIRaw storageTranscodequeueTranscodingworkersSegment storage(origin)Metadata DB+ cacheViewerCDN edges1. init upload2. upload parts3. uploaded4. segments5. status = readymetadatasegmentsmiss
Upload path (top) and watch path (bottom): the CDN serves the video bytes

6. Deep dives

Uploading reliably

  • Multipart, resumable uploads directly to object storage with pre-signed URLs: parallel parts, retry only failed parts, resume after network drops.
  • The metadata row starts as status = uploading. Completion triggers an event onto the transcode queue.
  • Validate early: file type, size limits, malware scan.

The transcoding pipeline

Raw video is huge and comes in random formats. Transcoding produces:

  • Multiple resolutions and bitrates (240p … 1080p, 4K) and codecs (H.264 for compatibility, VP9/AV1 for efficiency).
  • Segments of 2–6 seconds per rendition, plus manifests listing them.
  • Thumbnails, preview sprites and audio tracks.

Make it fast and robust:

  • Split the video into chunks (by keyframes) and transcode them in parallel on many workers, then stitch. A 1-hour video finishes in minutes instead of hours.
  • Model it as a DAG: split → transcode each rendition → package → thumbnails → publish. Each task is retried independently, and workers are idempotent (writing the same output key twice is harmless).
  • Prioritise: popular creators or short videos first; spare capacity for re-encoding old videos into better codecs.

Adaptive bitrate streaming (HLS / DASH)

Video playerMaster manifest1080p playlist(6 Mbps)480p playlist(1.5 Mbps)240p playlist(0.4 Mbps)1. fetch2. segments
The player picks a rendition per segment based on measured bandwidth
  • The master manifest lists renditions; each rendition’s playlist lists its segments.
  • The player measures throughput and buffer level and picks the best rendition for each segment, dropping quality instead of stalling when the network dips.
  • Short segments make switching responsive; longer segments are more efficient. 2–6 s is the usual balance.
  • Fast start: begin at a low bitrate, then climb.

CDN strategy

  • Segments are immutable, so they get Cache-Control: max-age=31536000, immutable and cache perfectly.
  • Popular videos are cached at (or pushed to) edges; the long tail is served from regional caches or origin shields, then the origin. Viewing follows a power law, so a small set of videos is most of the traffic.
  • Some platforms place cache appliances inside ISPs’ networks (Netflix Open Connect, Google Global Cache) to cut transit costs.

View counts

Counting every view synchronously in a database would be a hot-key nightmare for viral videos. Instead:

  • Clients send view events (batched); events go to a stream (Kafka).
  • A stream processor aggregates counts per video per minute and updates a counter store.
  • Displayed counts are eventually consistent (a few seconds behind). Dedupe and fraud filtering happen in the pipeline.

Storage cost

PB-scale storage needs lifecycle tiers: keep popular renditions hot; move rarely watched videos’ renditions to cheaper storage; even delete rarely used renditions and re-transcode on demand.

7. Bottlenecks and failure modes

What fails Impact Mitigation
Transcoding worker crash A chunk isn’t encoded Retry the chunk (idempotent tasks); queue redelivery
Transcode backlog Uploads take longer to go live Autoscale workers; prioritise; show “processing”
CDN edge outage Viewers in a region affected Multi-CDN, DNS steering to healthy edges, origin shield
Origin overload from misses Slow starts for long-tail videos Origin shield, regional caches, request collapsing
Viral video Huge, sudden demand Already cached at edges; pre-warm for scheduled premieres

8. Wrap-up

A strong answer separates the upload pipeline (direct-to-storage, queue, parallel DAG transcoding) from the watch path (HLS/DASH segments from a CDN), uses estimates to show the CDN carries the load, and handles view counting asynchronously.

Likely follow-ups: How would you add live streaming (lower-latency segments, ingest servers, no full transcode ahead of time)? How would you do recommendations? How do you handle copyright detection (fingerprinting in the pipeline)?

Test yourself

Answer in your head, then click a card to check. All cards are in the Anki deck.