LearnHLDDesign a notification service

Design a notification service

A marketing team schedules a campaign to 40 million users for 9am. At 9:00:02 the queue is full of campaign messages. At 9:00:03 someone’s one time password is behind 40 million of them, and by the time it arrives the login screen has already timed out.

Nobody wrote a bug. Every service worked. The design just never said which message mattered more, and that single omission is what this question is about.

Step 1: Understand the problem

“Send a notification” sounds like one API call. It is really a fan out problem, a scheduling problem and a third party integration problem wearing one name.

You askThey sayWhat it settles
Which channels: push, SMS, email, in app?All four, and a user may get the same message on more than one.A channel worker per channel, because a push token, a phone number and a mailbox fail in completely different ways and retry on completely different schedules.
How many notifications a day, and how bursty?About a billion, with campaign spikes ten times the average inside a minute.A queue between accepting and sending, and enough capacity planning to know the spike is scheduled rather than random.
Are all notifications equally urgent?No. A one time password is not a discount code.The most important answer on the page. Separate lanes with separate capacity, otherwise a campaign delays a login, which is the failure in the opening paragraph.
Can a user receive the same notification twice?Badly wrong for money and codes, merely annoying for news.An idempotency key from the caller and a dedupe window per user per message. At least once delivery plus a dedupe key is the honest answer, not exactly once.
Do users control what they receive?Yes. Per category opt outs, quiet hours, and a daily cap.A preference check before anything is enqueued, and a scheduler that can hold a message until morning instead of dropping it.
What do we owe the sending team afterwards?Whether it was delivered, and whether it was opened.A delivery event pipeline, and the awkward truth that "sent" is the only thing you actually control. The providers tell you the rest, late and incompletely.

What you are building, and what you cut

In scope
  • Accept a notification request from any internal service. One API, one contract, so no team is writing its own APNs client.
  • Fan out to the channels a user has enabled. Push, SMS, email and in app, in that cost order.
  • Respect preferences, quiet hours and rate caps. The cheapest notification is the one you decided not to send.
  • Retry failures and report delivery. Providers fail constantly and none of them fail the same way.
The numbers you commit to
  • A billion notifications a day, with 10x campaign spikes.
  • An urgent notification leaves the system in under 5 seconds at p99, always.
  • No duplicate for a given idempotency key inside 24 hours.
  • A provider outage delays notifications rather than losing them.
Cut, and say so out loud
  • Writing the copy, and deciding who to target. That is a campaign tool that calls this service.
  • In app inbox storage and read state. Similar words, different system, and it is a feed problem.
  • Delivery of the push payload itself. APNs and FCM own the last mile and you cannot see inside it.
  • Ranking or bundling by machine learning. Worth naming as the layer above the rate cap.

Back of the envelope

Send rate and the spike
500 million
2
10x
Notifications a day500M x 2 = 1.0 billion
Average send rate1,000,000,000 / 86,400 = 11,574 per second
Peak send rate11,574 x 10 = 115,741 per second
Device tokens stored500M x 2.4 devices = 144 GB
Urgent share of trafficabout 2%
115,741 per second
at peak, and 98% of it could wait an hour without anybody noticing

The gap between 11,574 and 115,741 is the entire capacity argument. You do not build for the peak of bulk traffic, you build a lane that lets the 2% that matters overtake it. Sizing the whole system for the campaign spike costs several times more and still does not fix the ordering problem.

The number people get wrong

Candidates size the fleet for peak send rate and stop. The real constraint is the provider. APNs will happily take a lot of traffic, most SMS gateways will not, and an email provider has a per second cap written into your contract. Your sender fleet can be twice as fast as your provider allows and it buys you nothing except a bigger queue.

Step 2: Propose the high level design

The API

One endpoint does almost all the work. The interesting decisions are all in the request body.

POST/v1/notifications
{
  "userId": 4471,
  "template": "otp_login",
  "params": { "code": "402913" },
  "priority": "urgent",
  "idempotencyKey": "otp:4471:1756112400"
}
returns 202 { "notificationId": "n_8f21", "status": "accepted" }
Why: 202 and not 200, because nothing has been sent yet and pretending otherwise makes every caller write a bug. The caller sends a template name and parameters rather than rendered text, so localisation, copy changes and channel specific formatting stay inside this service.
POST/v1/notifications:batch
{
  "template": "weekly_digest",
  "segmentId": "seg_9931",
  "priority": "bulk",
  "notBefore": "2026-09-11T09:00:00+05:30"
}
returns 202 { "campaignId": "c_1180", "estimated": 41200000 }
Why: A campaign is one request, not 40 million. Expanding a segment inside the service means the expansion happens at a rate you control, and cancelling a campaign that is going wrong is one call rather than a support incident.
PUT/v1/users/{id}/preferences
{
  "marketing": { "push": false, "email": true },
  "quietHours": { "start": "22:00", "end": "08:00", "tz": "Asia/Kolkata" },
  "dailyCap": 5
}
returns 200
Why: Preferences live here, not in each sending team's database. The moment a second team keeps its own opt out list, someone who unsubscribed gets a message and you are explaining it to a regulator.
GET/v1/notifications/{id}
returns 200 { status: "delivered", channel: "push", attempts: 2, deliveredAt: "..." }
Why: Sending teams will ask "did it arrive". Give them an answer or they will build their own tracking, badly. Note that status can stay "sent" forever, because some providers never confirm anything.

The data model

Three tables, and the interesting one is the smallest.

notificationswide column store, partitioned by user_id
user_idbigintPKPartition key. Everything about one person lives together, which is how the daily cap gets counted cheaply.
notification_iduuidPKClustering key, time ordered, so a user history read is one scan.
idempotency_keyvarchar(128)UQUnique per user for 24 hours. This one constraint is the entire duplicate defence.
templatevarchar(64)Not the rendered body. Storing rendered text means a copy fix cannot be applied retroactively and multiplies your storage.
priorityenumIDXurgent, standard, bulk. Chooses the queue, and nothing else in the system needs to know.
stateenumaccepted, suppressed, queued, sent, delivered, failed. Suppressed is a real outcome and gets counted.
attemptsintRetry count. Past the limit it goes to the dead letter queue with the last provider error attached.
Sample row
4471 | n_8f21... | otp:4471:1756112400 | otp_login | urgent | delivered | 1
Device tokens live in their own table keyed by user, because they change constantly and are read on every send. Preferences live in a third, cached aggressively, because they are read on every send and written almost never.

The whole system on one whiteboard

Figure 1. The three boxes above and below the queue are where notifications get stopped: preferences say no, the throttle says not this many, the scheduler says not right now. Everything that survives all three is worth paying a provider for.

Walking Figure 1:

  1. Any internal service posts a notification. It gets a 202 immediately, because nothing downstream is allowed to make a checkout page slow.
  2. The API renders the template for each channel. Callers never send rendered text, which is what lets you fix a typo in one place for every language.
  3. Preferences are checked before the message is enqueued. A suppressed notification is recorded as suppressed, not silently dropped, because “why didn’t my user get it” is the most common support question you will get.
  4. What survives is enqueued onto the lane matching its priority. Three lanes, three sets of workers, and nothing in the bulk lane can delay the urgent one.
  5. Dedupe and rate capping happen just before sending, not at accept time, because the cap is about what a person actually received.
  6. Workers pull at a rate the provider will tolerate, not at the rate the queue offers.
  7. They call the providers, record whatever callbacks come back, and hand permanent failures to a dead letter queue with the provider error attached.
Your answer

Step 4 splits into three lanes. If you had only one queue and doubled the number of workers instead, what exactly would still be broken?

Step 3: Design deep dive

Priority is a lane, not a number

The instinct is to put a priority field on the message and sort the queue by it. This does not work, for a reason worth being able to say precisely. A queue you have to sort is a queue you have to read, and a 40 million message backlog is not something you read inside a few milliseconds.

So separate the storage, not just the label. Three topics, three consumer groups, three independent capacity pools. Urgent is provisioned for its peak, which is small, and it is never behind anything. Bulk is provisioned for cost and is expected to lag.

1/6 The campaign arrives first, as one request. It has not become 40 million messages yet, and where that expansion happens is a real design decision.
Figure 2. One OTP and one campaign entering at the same moment. The lanes never touch, so the OTP overtakes 40 million messages without anything having to be sorted.
The follow up you will get

“What if urgent traffic itself spikes?” Then you are in trouble, and the honest answer is that urgent traffic is bounded by real human events, logins and payments, which do not spike tenfold in a minute. If a team starts marking marketing messages urgent, that is not a capacity problem, it is a governance one: enforce the priority per template and per caller, and put the urgent volume per team on a dashboard someone reviews.

Delivering at least once without annoying anybody

Every provider on the list can accept a request, fail to respond, and deliver anyway. That means retries are mandatory and duplicates are possible, so exactly once is not available. What you can build is at least once delivery with a deduplication window, which produces the same user experience for a fraction of the effort.

Three mechanisms, and you should name all three:

An idempotency key from the caller. For an OTP that key is naturally otp:{userId}:{minute}, so a retried login cannot produce a second message. Make it unique per user for 24 hours and enforce it in the database, not in a cache, because a cache eviction should never cost a duplicate.

A dedupe check in the worker, just before sending. This catches the notification that was enqueued twice by two different code paths, which happens more often than anyone expects once several teams are calling you.

The provider’s own idempotency token, where one exists. APNs has a collapse id, most SMS gateways have nothing, and email has nothing at all. So this third mechanism is partial, and you should say so rather than implying it covers everything.

The retry that makes it worse

Retrying immediately on a provider timeout is how a slow provider becomes a dead one. Exponential backoff with jitter, a cap on attempts, and a circuit breaker per provider, so that when APNs is degraded you stop hammering it and start queueing instead. Without the breaker, your retries are indistinguishable from an attack on your own vendor, and they will rate limit you at exactly the wrong moment.

The cap that stops people uninstalling the app

The most expensive failure in a notification system is not a dropped message. It is sending enough messages that somebody turns notifications off, because that user is now unreachable forever, including for the ones that matter.

So the service enforces limits the sending teams do not control:

A daily cap per user per category, counted in the same partition as the user’s notifications so the check is one read. A quiet hours window in the user’s own timezone, where a bulk notification is held until morning rather than dropped, and an urgent one goes anyway. And a collapse rule, where five notifications of the same category inside an hour become one that says five, which is the only mechanism here that reduces volume rather than moving it.

Say out loud that this belongs in the platform. Every team believes their notification is the important one, and a limit that each team implements for itself is not a limit.

Break it

Send pressure
12k/s
12k/s120k/s bulkprovider throttlesprovider downrepaired
Healthy. Average day. Queue depth is near zero on every lane, workers are half idle, and the providers are not noticing you at all.

The lesson worth taking from the last two states: in a system whose whole job is calling other people’s services, the circuit breaker is not a refinement. It is the difference between one channel degrading and everything stopping.

Trade-offs

ChoiceWhat you gainWhat you payPick it when
Separate queues per priorityUrgent messages are never behind bulk ones, with no sorting and no preemption.Three sets of consumers to run and monitor, and someone has to police who gets to say urgent.Always, the moment two classes of notification share a system. This is the cheapest correctness you will buy all day.
At least once with a dedupe keySurvives every provider timeout and every worker crash without losing messages.A duplicate is possible in the window where a provider delivered but never told you.Every time, because exactly once across a third party API does not exist and pretending it does hides the retry bug.
Expanding a campaign inside the serviceOne request instead of 40 million, expansion at a rate you control, and cancellation is one call.The service now owns segment expansion, which is state and work that feels like somebody else's job.Any campaign above about a hundred thousand recipients. Below that, let the caller loop.
Preferences and caps in the platformOne opt out that actually works, and a limit no individual team can talk its way around.A hot read on every send, and a service that has to be up for anything to be sent at all.Always. The regulatory version of this argument ends the discussion faster than the engineering one.

Interview replay

Interviewer
A service wants to notify a user. Take me through what happens.
Opening move. They want to see whether you make the caller wait.
You
The caller posts a template name, parameters and an idempotency key, and gets a 202 straight away. We render, check preferences, and enqueue onto the lane for that priority. Workers pull from the lane, check a dedupe set, look up device tokens and call the provider. Nothing in that path blocks the caller, because the caller is usually a checkout or a login and cannot afford to wait for APNs.
Names the 202 and the reason for it in the first answer. That is the difference between having used a notification service and having drawn one.
Interviewer
A campaign to 40 million users is in the queue and a one time password arrives. What happens?
The real question. Everything before was warm up.
You
They are in different queues, so the OTP is not behind anything. Sorting one queue by priority does not work at that depth, since you would have to read the backlog to sort it. Separate topics with separate consumer groups means the urgent lane is nearly always empty and its latency is independent of how much bulk traffic exists.
Rejects the obvious answer and explains why, rather than just asserting the better one.
Interviewer
Can you guarantee a user never gets the same notification twice?
A trap. The confident yes is wrong.
You
No, and I would not claim it. A provider can accept a request, time out on us, and deliver anyway. What I can do is at least once delivery with an idempotency key unique per user for 24 hours, enforced in the database, plus a dedupe check right before sending. That makes duplicates rare rather than impossible, and for the cases where a duplicate is genuinely unacceptable, like a payment confirmation, I would make the key deterministic from the payment id.
Refusing to overclaim, then giving the mechanism that gets closest. Much stronger than a yes.
Interviewer
APNs starts timing out. What does your system do?
Checking whether you have thought about being a client of something unreliable.
You
A circuit breaker per provider opens after a few failures, so workers stop blocking on socket timeouts and the push lane parks in the queue instead. Retries use exponential backoff with jitter and a cap, then the dead letter queue. The thing I would watch is queue age rather than queue depth, because depth looks alarming during a normal campaign and age tells you whether anything is actually stuck.
The queue age point is the one that sounds like operating experience, because it is.
Interviewer
How would you decide the daily cap per user?
Open ended. They are checking for an invented number.
You
I would not pick it analytically. I would measure the relationship between notifications received in a week and the rate at which people disable notifications entirely, then set the cap below wherever that curve turns. I have not run that experiment, so I would want the data before committing to a number, and I would expect it to differ a lot by category and by market.
Names the metric that actually matters, which is opt out rate rather than open rate, and admits the limit of what they know.

Checkpoint

Checkpoint

1. Why does a single queue with a priority field fail at 40 million messages?

2. A provider accepts a send, then times out before responding. What is the correct behaviour?

3. Same system, but now the product wants notifications ranked and bundled by relevance. What changes first?

Worth memorising
  • A billion a day is 12,000 a second average and 120,000 in a campaign minute. Build for the gap, not the peak.
  • Urgent traffic is about 2% of the total. That 2% is what the separate lane exists to protect.
  • Three lanes, three consumer groups. Sorting one queue that is 40 million deep is not a thing you do in milliseconds.
  • At least once, with a 24 hour dedupe key. Exactly once across a third party API does not exist.
Say this in 60 seconds

A notification service is mostly a fan out and a scheduling problem, plus being a well behaved client of providers you do not control. Internal services post a template name, parameters and an idempotency key and get a 202 immediately, because nothing downstream may block a login or a checkout. We render, check preferences and quiet hours, and enqueue onto one of three lanes by priority. The lanes are separate topics with separate consumers, which is what lets a one time password overtake a 40 million message campaign without anything being sorted. Workers pull at a rate the provider tolerates, dedupe just before sending, and retry with backoff behind a circuit breaker per provider, because the failure that actually happens is a vendor going slow and taking every worker with it. Delivery is at least once with a dedupe window, since exactly once across a third party API does not exist. And the platform, not the sending teams, owns the daily cap, because the most expensive failure here is somebody turning notifications off entirely.

IndGeek provides solutions in the software field, and is a hub for ultimate Tech Knowledge.