LearnHLDDesign a video streaming service

Design a video streaming service

A phone on a train plays a video at 1080p, goes through a tunnel, drops to 480p for eight seconds, and comes back up. Nobody notices. That is the product.

The interesting thing about this question is that almost none of the hard work happens while somebody is watching. Playback is a static file download from a cache near the user, and it is meant to be boring. Everything expensive happens once, when the video is uploaded, and everything clever happens on the client. Candidates who spend the interview designing the streaming path are designing the easy half.

Step 1: Understand the problem

The first question separates this into two completely different systems, so ask it first.

You askThey sayWhat it settles
On demand video, or live?On demand.Everything. On demand means transcode once and serve a static file forever. Live means transcoding in the request path with a latency budget, which is a different page and you should say so.
How many hours are uploaded, and how many watched?Thousands of hours uploaded a day. Hundreds of millions of hours watched.A ratio of about one to a hundred thousand. Upload is a batch problem you can queue and watching is a bandwidth problem you cannot, so they get completely separate infrastructure.
What devices and what networks?Everything from a 3G phone to a 4K television.Multiple renditions of every video, and adaptive switching between them mid playback. It is also why a single file per video is not an option.
How soon after upload must a video be watchable?Minutes, not hours.Transcoding is split into chunks and done in parallel, not as one long job per file. A two hour film transcoded serially is a two hour wait.
How exact does the view count have to be?Roughly right, quickly. Exact eventually, for payouts.A streaming counter for the display number and a batch recompute for the number creators are paid on. The same shape as the ad click aggregator, and worth saying so.
Do we need DRM and geo restrictions?Signed URLs and region blocking, yes. Full DRM, assume not.Short lived signed URLs at the CDN edge rather than an authorisation call per segment. An auth check on every six second chunk would be the busiest endpoint you own.

What you are building, and what you cut

In scope
  • Upload a video and make it watchable in minutes. Chunked upload, resumable, then parallel transcoding.
  • Transcode into several renditions. From 240p to 4K, plus multiple codecs.
  • Stream adaptively to any device. The client picks the rendition, segment by segment.
  • Count views. Approximately now, exactly later.
The numbers you commit to
  • Playback starts in under 2 seconds at p95.
  • Rebuffering under 0.5% of watch time.
  • A video is watchable within 5 minutes of upload finishing.
  • Tens of terabits a second of egress at peak, almost all of it from cache.
Cut, and say so out loud
  • Recommendations and the home page. It is a ranking system that happens to point at videos.
  • Live streaming, which puts transcoding on the latency path and needs its own design.
  • Comments, subscriptions and the social layer. A feed problem, covered elsewhere.
  • Content moderation and copyright matching. Enormous, and it runs beside the pipeline rather than inside it.

Back of the envelope

Where the money goes
10 thousand
500 million
7
Source video in per day10k hours x 3.6 GB = 36 TB
Stored after transcodingsource x 4.9 = 175 TB a day
Average egress500M hours at 2.5 Mbps = 416,667 Gbps
Peak egressaverage x 2.2 = 916,667 Gbps
Transcode CPU per day10k x 7 renditions = 28,000 core hours
916.7 Tbps
at peak, which is the number that decides whether you have a CDN strategy or a hobby

Storage grows by 175TB a day and never shrinks, but it is cheap and boring. The number that dominates the bill is egress. Everything in this design that looks like an optimisation is really an attempt to serve that 916.7 Tbps from a cache close to the user rather than from your own machines.

The number people get wrong

Candidates compute storage, find a large number, and design around it. Storage is the cheapest part of this system by an order of magnitude. Egress is the bill, and cache hit ratio at the edge is the single metric that moves it. A 95% hit ratio and a 99% hit ratio differ by five times the origin traffic, which is the difference between a normal origin fleet and a heroic one.

Step 2: Propose the high level design

The API

POST/v1/uploads
{ "filename": "final_cut_v3.mov", "sizeBytes": 8140000000, "sha256": "..." }
returns 201 { uploadId, partUrls: [...], partSizeBytes: 8388608 }
Why: The server hands back pre signed URLs and the client uploads parts straight to object storage. Proxying eight gigabytes through your API servers wastes bandwidth twice and turns every flaky upload into a held connection you are paying for.
PUT{presigned part url}
returns 200 with an ETag per part
Why: Resumable by construction. A phone that loses signal at part 340 of 970 resumes at 340, and a client that never returns leaves parts that a lifecycle rule deletes after a week.
GET/v1/videos/{id}/manifest.m3u8
returns 200, a text manifest listing renditions and segment URLs
Why: The manifest is the whole streaming API. It is a small text file, cached for seconds rather than hours, and it points at segment URLs cached for a year. Everything about adaptive playback is a client reading this file and choosing.
POST/v1/videos/{id}/heartbeat
{ "positionMs": 184000, "rendition": "720p", "bufferMs": 12000, "sessionId": "..." }
returns 204
Why: Sent every 30 seconds while playing. This is how you count a view honestly, and the buffer and rendition fields are what tell you your delivery is degrading before anybody complains on social media.

The data model

videosrelational for metadata, object storage for bytes
video_iduuidPKAssigned at upload start, before a single byte arrives.
stateenumIDXuploading, transcoding, ready, failed, blocked. The client polls this, and it is the only thing standing between a user and a broken player.
source_keyvarchar(256)Where the original lives. Kept forever, because every future codec means transcoding it again.
duration_msbigintFrom the probe step, not from the uploader.
renditionsjsonWhich resolutions and codecs actually completed. A video can be watchable at 480p while 4K is still encoding.
region_policyjsonWhere it may be played. Enforced at the edge, not in the player.
Sample row
a91f... | ready | src/a91f/original.mov | 7,384,000 | [240p, 480p, 720p, 1080p] | {block: [XX]}
Segments are not in a database. They are files in object storage, named deterministically as video/rendition/segment_00042.ts, so the manifest can list them without a lookup and the CDN can cache them by URL forever. Making segment locations computable rather than stored removes an entire tier from the hot path.

The whole system on one whiteboard

Figure 1. The upload pipeline across the top runs once per video and is allowed to take minutes. The playback path along the bottom runs a billion times and mostly never reaches you at all.

Walking Figure 1:

  1. The client uploads parts directly to object storage with pre signed URLs. Your API never touches the bytes, which is the single biggest saving in the upload path.
  2. Completion triggers the pipeline. The source is probed and split into chunks of about thirty seconds, cut on keyframe boundaries so each chunk is independently decodable.
  3. Each chunk becomes one job per rendition. A two hour film at seven renditions is a few thousand small jobs rather than one enormous one.
  4. Workers pick them up. This fleet is the largest compute cost in the company and it is a perfect fit for interruptible spot capacity, because a lost job is just a retried chunk.
  5. The packager assembles encoded chunks into playable segments and writes the manifests.
  6. Metadata flips to ready, per rendition, so 480p can be live while 4K is still going.
  7. Playback: the viewer fetches segments from the nearest edge. Almost all of this never reaches your infrastructure.
  8. The manifest comes from your API, because it carries signed URLs and region rules and is the only per user part of playback.
  9. Heartbeats feed the view counting pipeline, entirely off the delivery path.
Your answer

Step 4 says a lost transcode job is 'just a retried chunk'. What property of step 2 makes that true, and what would break if the splitter cut chunks at fixed byte offsets instead?

Step 3: Design deep dive

Transcoding is a fan out problem, not a video problem

The naive pipeline takes a two hour film and runs ffmpeg on it, seven times. That is many hours of wall clock time on one machine, no way to use idle capacity elsewhere, and a single crash at 90% loses everything.

The fix is to stop thinking about the file.

1/6 The trigger is the storage event, not an API call, so an upload that finishes while your service is deploying still gets processed.
Figure 2. One upload becomes a few thousand independent jobs, runs on cheap interruptible machines, and finishes in minutes. The only tricky part is where you are allowed to cut.

Two details worth saying unprompted. Transcode the low renditions first, because being watchable at 480p in ninety seconds is worth far more than being watchable at 1080p in six minutes. And keep the source forever: every new codec, and there is always a new codec, means running this whole pipeline again over the back catalogue.

The follow up you will get

“How would you handle a two second video and a four hour one in the same pipeline?” Chunking a two second file is pure overhead, so below some threshold you transcode it whole in one job. Above it, chunk. The threshold exists because the fixed cost of a job is roughly constant, so state it as a rule with a number attached rather than as a special case: below about a minute, one job, and above it, thirty second chunks.

The client decides, not the server

Adaptive bitrate sounds like a server feature and is almost entirely a client one. The server publishes a manifest describing what exists. The player measures its own throughput and buffer, and picks the next segment.

Who picks the quality
The player knows things the server cannot: the size of its buffer, the screen it is drawing on, whether the last segment arrived slowly, and whether the user is on a metered connection. Every segment is a static file so the server has no per session state at all, which is what lets a CDN serve almost everything. This is what HLS and DASH do and it is the right answer.

The consequence to state explicitly: because the client decides, every segment is an ordinary immutable file with a URL. That is why a CDN can hold 99% of your traffic, and that is why the egress bill is survivable. Statelessness is not an aesthetic preference here, it is the business model.

Two client behaviours are worth naming because they show you have debugged a player. Start at a low rendition and climb, since a fast first segment matters more than a beautiful one. And switch down aggressively but up conservatively, because an unnecessary downgrade is invisible and a rebuffer is not.

Getting the bytes near the user

At a few terabits a second, the only question that matters is what fraction is served from cache.

Segments are perfect cache objects. They are immutable, they are addressed by URL, they are a few megabytes each, and popular ones are extremely popular. Set a one year TTL and never invalidate: if a video changes, it gets new URLs.

The problem is the long tail. Most videos have almost no views, so the first viewer in a region always misses, and with hundreds of edge locations one unpopular video can produce hundreds of origin fetches for the same segment. That is what an origin shield is for: a middle tier that edges fetch through, turning three hundred misses into one.

What large services actually do

Netflix ships physical caching appliances into internet providers’ own data centres, so the popular catalogue is already inside the network your customer is connected to before they press play. YouTube does the same thing with its own edge nodes. The interview version of this insight is short and worth saying: at this scale, delivery stops being a software problem and becomes a question of where you put hardware.

Break it

Viewing load
normal evening
normal eveningnew releasecatalogue refreshedge site downrepaired
Healthy. 99% hit ratio at the edge. Origin sees a trickle, mostly the long tail, and the whole system is doing what it was built to do: almost nothing.

The catalogue refresh state is the one worth raising unprompted, because it is a self inflicted outage that looks like a routine deploy, and every candidate who has lived through one mentions it before being asked.

Trade-offs

ChoiceWhat you gainWhat you payPick it when
Chunked parallel transcodingMinutes instead of hours, and interruptible machines become usable because a lost job is one chunk.Cuts must land on keyframes, and quality across chunk boundaries needs care.Any file long enough that serial encoding is measured in hours. Below about a minute, do it in one job.
Client side adaptive bitrateEvery segment is an immutable file, so the CDN serves nearly everything and the server holds no session state.You give up central control of quality, and a badly written player degrades the experience with no server side fix.On demand video, always. Server side selection only makes sense when you must control quality centrally, and it costs you cacheability.
Pre signed direct uploadEight gigabyte files never pass through your API tier, and resumability is free.Less control over what arrives, so validation moves to a probe step after the fact.Any upload above a few megabytes. Proxying large files through application servers is a habit worth breaking.
Origin shield in front of storageHundreds of edge misses for the same object collapse into one origin fetch.Another tier to operate, and one more hop of latency on a cold read.Any large catalogue with many edge locations and a long tail, which is every video service.

Interview replay

Interviewer
Someone presses play. Walk me through it.
Opening. They are watching for whether you invent a streaming server.
You
The player fetches a manifest from our API, which is the only per user part because it carries signed URLs and region rules. The manifest lists renditions and segment URLs. From then on the player requests six second segments directly from the nearest CDN edge and picks the quality itself based on its buffer and measured throughput. Ninety nine percent of those requests never reach our infrastructure, because a segment is an immutable file with a stable URL.
Establishes statelessness immediately, which is the property everything else depends on.
Interviewer
Why not have the server choose the bitrate? It knows more about the network.
Testing whether you can defend the choice rather than recite it.
You
It knows less, actually. The player knows its buffer depth, its screen size, and how slowly the last segment arrived. But the real reason is cacheability: the moment the server chooses per viewer, every response is personalised and nothing can be cached at the edge. We would go from a 99% hit ratio to zero and multiply our egress bill by a hundred, to make a worse decision.
Gives the technical reason and then the economic one. The second is what makes it memorable.
Interviewer
A two hour film is uploaded. How long until someone can watch it?
Checking whether transcoding is a fan out problem in your head.
You
A couple of minutes for the first watchable rendition. We split the source on keyframes into roughly thirty second chunks, which for two hours is about 240 chunks, and each chunk times each rendition is an independent job. Thousands of small jobs across a large spot fleet finish in minutes rather than hours. I would encode the low renditions first and publish each rendition as it completes, so 480p is live while 4K is still running.
Names the chunk count and the publish-as-you-go behaviour. Both are concrete and both are what actually gets shipped.
Interviewer
Your CDN hit ratio drops from 99% to 80% overnight. What happened?
A debugging question in a design interview. They want a hypothesis, not a procedure.
You
Most likely something changed segment URLs, so every cache went cold at once. A codec rollout, a change to how we sign URLs, or a cache key that accidentally started including a query parameter. That last one is the classic: add an analytics parameter to segment URLs and every request becomes unique. I would check whether origin traffic rose uniformly across regions, which points at a URL change, or in one region, which points at an edge problem.
Gives a specific likely cause and a way to distinguish between two hypotheses. Much stronger than "I would look at the logs".
Interviewer
How many renditions would you generate?
Open ended, and expensive to get wrong in both directions.
You
I would not fix a number. Every rendition costs storage and transcode time on every video, including the ninety percent nobody watches, so the honest approach is to generate a small ladder for everything and the full ladder only once a video crosses a view threshold. What that threshold should be depends on the popularity distribution, which I would want to measure rather than guess. I know large services do per title encoding, choosing a ladder based on the complexity of the actual content, but I have not built that and I would treat it as a later optimisation.
Turns a number question into a policy, and names a technique they have heard of without pretending to have implemented it.

Checkpoint

Checkpoint

1. Why must the splitter cut chunks on keyframe boundaries?

2. What does client side bitrate selection actually buy the system?

3. Same design, but now it is live streaming with a five second delay target. What breaks first?

Worth memorising
  • Egress is the bill. Storage is cheap and transcoding is a one time cost per video.
  • A 95% hit ratio instead of 99% is five times the origin traffic. Hit ratio is the metric that moves money.
  • 7 renditions x 240 chunks for a two hour film is a few thousand independent jobs.
  • Six second immutable segments with a one year TTL. Anything that changes their URLs is a self inflicted outage.
Say this in 60 seconds

Video on demand is two systems. The upload pipeline runs once per video and is allowed to take minutes: the client uploads parts directly to object storage with pre signed URLs, a splitter probes the file and cuts it on keyframes into thirty second chunks, and each chunk times each rendition becomes an independent job on a large spot fleet. Low renditions encode first and each rendition publishes as it completes, so a two hour film is watchable at 480p in about two minutes. The playback path is deliberately boring. The player fetches a manifest, which is the only per user part, and then pulls immutable segments from the nearest CDN edge, choosing quality itself from its own buffer and throughput. That client side decision is what keeps every segment cacheable, and cacheability is the whole economic model, because egress at a few terabits a second is the bill and storage is not. An origin shield collapses hundreds of cold edge misses into one. The failure I would call out is anything that changes segment URLs across the catalogue at once, because it takes the hit ratio from ninety nine percent to zero and multiplies origin load by a hundred.

IndGeek provides solutions in the software field, and is a hub for ultimate Tech Knowledge.