Design a notification service
A marketing team schedules a campaign to 40 million users for 9am. At 9:00:02 the queue is full of campaign messages. At 9:00:03 someone’s one time password is behind 40 million of them, and by the time it arrives the login screen has already timed out.
Nobody wrote a bug. Every service worked. The design just never said which message mattered more, and that single omission is what this question is about.
Step 1: Understand the problem
“Send a notification” sounds like one API call. It is really a fan out problem, a scheduling problem and a third party integration problem wearing one name.
| You ask | They say | What it settles |
|---|---|---|
| Which channels: push, SMS, email, in app? | All four, and a user may get the same message on more than one. | A channel worker per channel, because a push token, a phone number and a mailbox fail in completely different ways and retry on completely different schedules. |
| How many notifications a day, and how bursty? | About a billion, with campaign spikes ten times the average inside a minute. | A queue between accepting and sending, and enough capacity planning to know the spike is scheduled rather than random. |
| Are all notifications equally urgent? | No. A one time password is not a discount code. | The most important answer on the page. Separate lanes with separate capacity, otherwise a campaign delays a login, which is the failure in the opening paragraph. |
| Can a user receive the same notification twice? | Badly wrong for money and codes, merely annoying for news. | An idempotency key from the caller and a dedupe window per user per message. At least once delivery plus a dedupe key is the honest answer, not exactly once. |
| Do users control what they receive? | Yes. Per category opt outs, quiet hours, and a daily cap. | A preference check before anything is enqueued, and a scheduler that can hold a message until morning instead of dropping it. |
| What do we owe the sending team afterwards? | Whether it was delivered, and whether it was opened. | A delivery event pipeline, and the awkward truth that "sent" is the only thing you actually control. The providers tell you the rest, late and incompletely. |
What you are building, and what you cut
- Accept a notification request from any internal service. One API, one contract, so no team is writing its own APNs client.
- Fan out to the channels a user has enabled. Push, SMS, email and in app, in that cost order.
- Respect preferences, quiet hours and rate caps. The cheapest notification is the one you decided not to send.
- Retry failures and report delivery. Providers fail constantly and none of them fail the same way.
- A billion notifications a day, with 10x campaign spikes.
- An urgent notification leaves the system in under 5 seconds at p99, always.
- No duplicate for a given idempotency key inside 24 hours.
- A provider outage delays notifications rather than losing them.
- Writing the copy, and deciding who to target. That is a campaign tool that calls this service.
- In app inbox storage and read state. Similar words, different system, and it is a feed problem.
- Delivery of the push payload itself. APNs and FCM own the last mile and you cannot see inside it.
- Ranking or bundling by machine learning. Worth naming as the layer above the rate cap.
Back of the envelope
The gap between 11,574 and 115,741 is the entire capacity argument. You do not build for the peak of bulk traffic, you build a lane that lets the 2% that matters overtake it. Sizing the whole system for the campaign spike costs several times more and still does not fix the ordering problem.
Candidates size the fleet for peak send rate and stop. The real constraint is the provider. APNs will happily take a lot of traffic, most SMS gateways will not, and an email provider has a per second cap written into your contract. Your sender fleet can be twice as fast as your provider allows and it buys you nothing except a bigger queue.
Step 2: Propose the high level design
The API
One endpoint does almost all the work. The interesting decisions are all in the request body.
{
"userId": 4471,
"template": "otp_login",
"params": { "code": "402913" },
"priority": "urgent",
"idempotencyKey": "otp:4471:1756112400"
}{
"template": "weekly_digest",
"segmentId": "seg_9931",
"priority": "bulk",
"notBefore": "2026-09-11T09:00:00+05:30"
}{
"marketing": { "push": false, "email": true },
"quietHours": { "start": "22:00", "end": "08:00", "tz": "Asia/Kolkata" },
"dailyCap": 5
}The data model
Three tables, and the interesting one is the smallest.
| user_id | bigint | PK | Partition key. Everything about one person lives together, which is how the daily cap gets counted cheaply. |
| notification_id | uuid | PK | Clustering key, time ordered, so a user history read is one scan. |
| idempotency_key | varchar(128) | UQ | Unique per user for 24 hours. This one constraint is the entire duplicate defence. |
| template | varchar(64) | Not the rendered body. Storing rendered text means a copy fix cannot be applied retroactively and multiplies your storage. | |
| priority | enum | IDX | urgent, standard, bulk. Chooses the queue, and nothing else in the system needs to know. |
| state | enum | accepted, suppressed, queued, sent, delivered, failed. Suppressed is a real outcome and gets counted. | |
| attempts | int | Retry count. Past the limit it goes to the dead letter queue with the last provider error attached. |
The whole system on one whiteboard
Walking Figure 1:
- Any internal service posts a notification. It gets a 202 immediately, because nothing downstream is allowed to make a checkout page slow.
- The API renders the template for each channel. Callers never send rendered text, which is what lets you fix a typo in one place for every language.
- Preferences are checked before the message is enqueued. A suppressed notification is recorded as suppressed, not silently dropped, because “why didn’t my user get it” is the most common support question you will get.
- What survives is enqueued onto the lane matching its priority. Three lanes, three sets of workers, and nothing in the bulk lane can delay the urgent one.
- Dedupe and rate capping happen just before sending, not at accept time, because the cap is about what a person actually received.
- Workers pull at a rate the provider will tolerate, not at the rate the queue offers.
- They call the providers, record whatever callbacks come back, and hand permanent failures to a dead letter queue with the provider error attached.
Step 4 splits into three lanes. If you had only one queue and doubled the number of workers instead, what exactly would still be broken?
Step 3: Design deep dive
Priority is a lane, not a number
The instinct is to put a priority field on the message and sort the queue by it. This does not work, for a reason worth being able to say precisely. A queue you have to sort is a queue you have to read, and a 40 million message backlog is not something you read inside a few milliseconds.
So separate the storage, not just the label. Three topics, three consumer groups, three independent capacity pools. Urgent is provisioned for its peak, which is small, and it is never behind anything. Bulk is provisioned for cost and is expected to lag.
“What if urgent traffic itself spikes?” Then you are in trouble, and the honest answer is that urgent traffic is bounded by real human events, logins and payments, which do not spike tenfold in a minute. If a team starts marking marketing messages urgent, that is not a capacity problem, it is a governance one: enforce the priority per template and per caller, and put the urgent volume per team on a dashboard someone reviews.
Delivering at least once without annoying anybody
Every provider on the list can accept a request, fail to respond, and deliver anyway. That means retries are mandatory and duplicates are possible, so exactly once is not available. What you can build is at least once delivery with a deduplication window, which produces the same user experience for a fraction of the effort.
Three mechanisms, and you should name all three:
An idempotency key from the caller. For an OTP that key is naturally
otp:{userId}:{minute}, so a retried login cannot produce a second message. Make it
unique per user for 24 hours and enforce it in the database, not in a cache, because a
cache eviction should never cost a duplicate.
A dedupe check in the worker, just before sending. This catches the notification that was enqueued twice by two different code paths, which happens more often than anyone expects once several teams are calling you.
The provider’s own idempotency token, where one exists. APNs has a collapse id, most SMS gateways have nothing, and email has nothing at all. So this third mechanism is partial, and you should say so rather than implying it covers everything.
Retrying immediately on a provider timeout is how a slow provider becomes a dead one. Exponential backoff with jitter, a cap on attempts, and a circuit breaker per provider, so that when APNs is degraded you stop hammering it and start queueing instead. Without the breaker, your retries are indistinguishable from an attack on your own vendor, and they will rate limit you at exactly the wrong moment.
The cap that stops people uninstalling the app
The most expensive failure in a notification system is not a dropped message. It is sending enough messages that somebody turns notifications off, because that user is now unreachable forever, including for the ones that matter.
So the service enforces limits the sending teams do not control:
A daily cap per user per category, counted in the same partition as the user’s notifications so the check is one read. A quiet hours window in the user’s own timezone, where a bulk notification is held until morning rather than dropped, and an urgent one goes anyway. And a collapse rule, where five notifications of the same category inside an hour become one that says five, which is the only mechanism here that reduces volume rather than moving it.
Say out loud that this belongs in the platform. Every team believes their notification is the important one, and a limit that each team implements for itself is not a limit.
Break it
The lesson worth taking from the last two states: in a system whose whole job is calling other people’s services, the circuit breaker is not a refinement. It is the difference between one channel degrading and everything stopping.
Trade-offs
| Choice | What you gain | What you pay | Pick it when |
|---|---|---|---|
| Separate queues per priority | Urgent messages are never behind bulk ones, with no sorting and no preemption. | Three sets of consumers to run and monitor, and someone has to police who gets to say urgent. | Always, the moment two classes of notification share a system. This is the cheapest correctness you will buy all day. |
| At least once with a dedupe key | Survives every provider timeout and every worker crash without losing messages. | A duplicate is possible in the window where a provider delivered but never told you. | Every time, because exactly once across a third party API does not exist and pretending it does hides the retry bug. |
| Expanding a campaign inside the service | One request instead of 40 million, expansion at a rate you control, and cancellation is one call. | The service now owns segment expansion, which is state and work that feels like somebody else's job. | Any campaign above about a hundred thousand recipients. Below that, let the caller loop. |
| Preferences and caps in the platform | One opt out that actually works, and a limit no individual team can talk its way around. | A hot read on every send, and a service that has to be up for anything to be sent at all. | Always. The regulatory version of this argument ends the discussion faster than the engineering one. |
Interview replay
Checkpoint
1. Why does a single queue with a priority field fail at 40 million messages?
2. A provider accepts a send, then times out before responding. What is the correct behaviour?
3. Same system, but now the product wants notifications ranked and bundled by relevance. What changes first?
- A billion a day is 12,000 a second average and 120,000 in a campaign minute. Build for the gap, not the peak.
- Urgent traffic is about 2% of the total. That 2% is what the separate lane exists to protect.
- Three lanes, three consumer groups. Sorting one queue that is 40 million deep is not a thing you do in milliseconds.
- At least once, with a 24 hour dedupe key. Exactly once across a third party API does not exist.
A notification service is mostly a fan out and a scheduling problem, plus being a well behaved client of providers you do not control. Internal services post a template name, parameters and an idempotency key and get a 202 immediately, because nothing downstream may block a login or a checkout. We render, check preferences and quiet hours, and enqueue onto one of three lanes by priority. The lanes are separate topics with separate consumers, which is what lets a one time password overtake a 40 million message campaign without anything being sorted. Workers pull at a rate the provider tolerates, dedupe just before sending, and retry with backoff behind a circuit breaker per provider, because the failure that actually happens is a vendor going slow and taking every worker with it. Delivery is at least once with a dedupe window, since exactly once across a third party API does not exist. And the platform, not the sending teams, owns the daily cap, because the most expensive failure here is somebody turning notifications off entirely.
