A notification system turns domain events into messages on email, push, SMS and in-app channels. The design centers on decoupling producers from delivery, respecting per-user preferences, and handling provider failures without losing or duplicating messages.
Before you start
You should understand queues, idempotency and fan-out. This article is conceptual with small sketches.
Step-by-step walkthrough
Step 1: Decouple producers from delivery
Producers publish an event to a queue instead of calling a provider directly, so a slow or failing provider does not block the request that triggered the notification. A worker consumes events and decides what to send, which lets you scale delivery independently.
Step 2: Apply preferences and deduplication
Before sending, check the user’s channel preferences and quiet hours, and deduplicate so the same event does not send twice, for example by a key of eventId + userId + channel. This step is where most product bugs come from, so it deserves explicit storage.
Step 3: Retry with failover and a digest option
Transient provider failures need retries with backoff and a fallback provider or channel. For high-volume updates, batch into a digest instead of sending one message per event, which reduces noise and provider load.
Worked scenario
The worker checks preference and a dedupe key before sending.
event(userId=42, type=order.shipped) -> queue
worker:
if channel disabled by preference: drop
if dedupeKey seen: drop
send via provider, on failure retry with backoffWalk through the example
The event is queued rather than sent inline, so the shipping service is not coupled to the email provider. The worker drops disabled channels and already-seen keys, then sends with retries. Each decision is a filter that reduces duplicate or unwanted messages.
Common mistake
Sending synchronously from the domain service, which couples latency and failures to the provider. Another is having no deduplication, so a reprocessed event spams the user, or ignoring preferences entirely.
Verify the behavior
Replay an event and confirm the dedupe key prevents a second send. Disable a channel and confirm no message goes out on it. Fail the primary provider and confirm the retry or fallback delivers, and that a digest batches multiple events into one message.
Interview exercise
Why queue notification events instead of sending from the service that raised them?
Answer and reasoning
Queuing decouples the producer from delivery, so a slow provider does not slow the domain request and a provider outage does not fail the business operation. It also enables retries, ordering and scaling independent of the producer, and it gives a durable record of what should be sent. The cost is eventual delivery, which is acceptable for notifications.
Continue learning
Compare fan-out in Feed fan-out and delivery in SQS delivery. Read the AWS messaging guidance and try the System design interview questions.