LearnHLDBack of the envelope math

Back of the envelope math

About eight minutes into a system design round, the interviewer asks how much traffic this thing has to take. What they are checking is not your arithmetic. They are checking whether the design you are about to draw is a response to a number or a response to a blog post you read.

There are only three numbers worth computing, and each one buys you a specific decision.

Your answer

A service has 10 million daily active users who each make 20 requests a day. Roughly what is the average requests per second, and what would you guess the peak is?

The only trick you need

Seconds in a day is 86,400. Round it to 100,000, which is 10^5. Now every traffic calculation is subtraction of exponents, and you can do it out loud without a whiteboard.

10 million users times 20 requests is 200 million requests a day. That is 2 x 10^8. Divide by 10^5 and you get 2 x 10^3, so 2,000 requests a second. You just rounded 86,400 up by 16%, which makes your answer 16% low, and nobody in the room cares. What they care about is that you got to a number in nine seconds.

Round hard, then say you rounded

Say “call it 100,000 seconds in a day” out loud. It signals that you know the real number and chose to drop it. Saying “86,400” and then reaching for a calculator signals the opposite.

Number one: requests per second

Traffic
10 million
20
3x
Requests per day10M x 20 = 200,000,000
Average per seconddaily / 86,400 = 2,315 rps
Peak per second2,315 x 3 = 6,944 rps
6,944 rps
peak, and this is the number you design for

Design for the peak, not the average. A system sized for 2,315 rps falls over every evening at eight.

The peak multiplier is the part people skip. Traffic is not flat. A consumer app in India does most of its work between 8pm and 11pm, so three times average is a reasonable starting guess and you should say why you picked it. A B2B tool concentrated in one working day is worse, closer to five. A system that everyone hits at the same instant, a ticket sale or a match starting, does not have a multiplier at all. It has a spike, and that is a different design problem.

Number two: storage

Storage is where candidates spend too long and get too little credit. It is a multiplication, and the only judgement in it is the size of one record.

Storage
50 million
500
5 years
Raw50M x 500B x 365 x 5 = 46 TB
With indexes, roughly 1.3x59 TB
With 3 replicas178 TB
178 TB
what you actually have to buy

The raw number is the one people quote and the replicated number is the one that shows up on the bill. Say both.

Two multipliers that almost nobody applies out loud, and both of them earn a nod: indexes add something like 30% on a normal relational schema, and replication multiplies everything by three. Fifty terabytes of raw data is two hundred terabytes of disk.

Number three: bandwidth

Peak requests per second times bytes per response. That is it. It matters in exactly two situations: when you are serving media, and when you are fanning one write out to many readers.

Ten thousand image requests a second at 200KB each is 2GB a second, or 16 gigabits. That is a real cost and a real reason to put a CDN in the drawing. Ten thousand JSON responses a second at 2KB each is 20MB a second, which is nothing, and computing it was a waste of forty seconds you could have spent on the data model.

The numbers you should not have to derive

Guess each one before you reveal it. The ones you get wrong by more than 10x are the ones worth writing down.

Latency, roughly, on modern hardware
Read 1MB sequentially from memory
Round trip inside one datacenter
Read 1MB sequentially from an SSD
Mumbai to Singapore round trip
Mumbai to Virginia round trip
Sizes, to two significant figures
A UUID stored as text
A typical JSON API response
A compressed photo from a phone
One minute of 1080p video
Rows a single well indexed relational primary can serve per second

What not to compute

ChoiceWhat you gainWhat you payPick it when
Peak requests per secondSizes the stateless tier and tells you whether you need a cache at all.Thirty seconds.Always. This is the number the rest of the round hangs off.
Total storage over the retention windowDecides whether one database is enough or you are in sharding territory.Thirty seconds, and a guess at record size.Always, but keep it to one line. Nobody has ever failed a round for a wrong record size.
BandwidthJustifies a CDN and exposes fan out costs.Twenty seconds.Media, or one write read by many. Skip it for a JSON API and say why you skipped it.
Memory for the cacheTurns "add a cache" into "add 40GB of cache", which is a different sentence.A guess at the hot fraction, usually 20%.Read heavy systems, once you have already decided a cache belongs.

Things not on that list: CPU cores, thread counts, connection pool sizes, exact disk IOPS. They depend on details nobody has in a 45 minute conversation, and reaching for them reads as someone filling time rather than someone making a decision.

Interview replay

This is the one page where the replay is not about a design. It is about doing arithmetic out loud without losing the room.

Interviewer
Design a photo sharing app. Assume 100 million daily users.
The number is bait. They want to see whether you do anything with it or start drawing boxes.
You
Let me get two numbers on the board first. A hundred million daily users, and say each one uploads a photo every other day, so 50 million uploads a day. That is 100,000 seconds in a day, near enough, so about 500 uploads a second average and maybe 1,500 at peak. Reads are the other side: if each user opens the app ten times and sees twenty photos, that is 20 billion photo views a day, so 200,000 a second. Four hundred to one.
Rounds 86,400 to 100,000 without apologising, and lands on the ratio rather than on either number. The ratio is what decides the design.
Interviewer
Where did the every other day come from?
Checking whether the assumption was deliberate or invented.
You
I made it up, and I would want to check it. It is the number I am least sure about, so it is worth saying out loud: if it is one photo a day instead, uploads double to 3,000 a second at peak, which still does not change the shape of the design because reads dominate either way. If it were fifty photos a day, that would change things, and I would want data before assuming it.
Naming the shakiest assumption and then testing whether the answer is sensitive to it is the single most scoreable move in estimation.
Interviewer
How much storage after five years?
Straightforward, and the trap is precision.
You
Fifty million uploads a day at, say, 2 megabytes after compression, is 100 terabytes a day. Times 365 is about 36 petabytes a year, so roughly 180 petabytes over five years, and with replication call it half an exabyte. I would not treat that as a precise figure. What it tells me is that this is object storage with a lifecycle policy, not a database question, and that cold storage tiering is worth real money here.
Ends on what the number means rather than on the number. A figure with no consequence attached is a figure nobody remembers.
Interviewer
And bandwidth?
The number most candidates skip, and it is often the one that costs the most.
You
200,000 photo views a second at 200 kilobytes for a feed sized image is about 40 gigabytes a second, which is 320 gigabits a second. That is the number that says this is a CDN problem before it is a server problem, because serving that from origin is a completely different budget from serving one percent of it.
Converting bytes to bits before quoting a network number, and naming the consequence, is what makes this sound rehearsed in a good way.
Interviewer
You are estimating quite fast. How confident are you?
An invitation to either overclaim or show judgement.
You
Confident within a factor of two or three, which is all these are for. I am using them to choose between architectures, not to size a purchase order. If two designs are only distinguishable by a number I am unsure about, I would say so and pick based on something else, like operational cost or how each one fails.
States the accuracy the exercise actually needs. Claiming precision on a napkin estimate is a much worse answer than admitting the range.

Checkpoint

Checkpoint

1. 10 million daily users at 20 requests each. What is the average requests per second?

2. You compute 50TB of raw data. What number should you actually quote to the interviewer?

3. Your API returns 2KB JSON responses at a peak of 8,000 requests a second. How much time should you spend on the bandwidth calculation?

Worth memorising
  • 100,000 seconds in a day. Round it, say out loud that you rounded, and move on.
  • 1 million a day is about 12 a second. 1 billion a day is about 12,000 a second. Two anchors cover most questions.
  • Memory 100ns, SSD 100us, same datacentre round trip 500us, cross continent 150ms. Four latencies, four orders of magnitude.
  • A factor of two is close enough. These numbers choose between architectures, not purchase orders.
Say this in 60 seconds

I would start with peak requests per second, because that sizes the stateless tier and tells me whether I need a cache. Daily active users times requests per user, divided by a hundred thousand seconds, times a peak multiplier of about three for a consumer app. Then storage: records per day times record size times retention, and I would quote the number with indexes and three replicas because that is what you actually buy. Bandwidth only if we are serving media or fanning out. I would not compute CPU or connection pools in a design round, since those depend on details we do not have here.

IndGeek provides solutions in the software field, and is a hub for ultimate Tech Knowledge.