Ch. 19 · System Design

System Design: Notification System

Ingest events, respect preferences, deduplicate and deliver across channels with retries and provider failover.

~2 min readadvancedupdated Oct 5, 2026

A notification system turns domain events into messages on email, push, SMS and in-app channels. The design centers on decoupling producers from delivery, respecting per-user preferences, and handling provider failures without losing or duplicating messages.

Before you start

You should understand queues, idempotency and fan-out. This article is conceptual with small sketches.

Step-by-step walkthrough

Step 1: Decouple producers from delivery

Producers publish an event to a queue instead of calling a provider directly, so a slow or failing provider does not block the request that triggered the notification. A worker consumes events and decides what to send, which lets you scale delivery independently.

Step 2: Apply preferences and deduplication

Before sending, check the user’s channel preferences and quiet hours, and deduplicate so the same event does not send twice, for example by a key of eventId + userId + channel. This step is where most product bugs come from, so it deserves explicit storage.

Step 3: Retry with failover and a digest option

Transient provider failures need retries with backoff and a fallback provider or channel. For high-volume updates, batch into a digest instead of sending one message per event, which reduces noise and provider load.

Worked scenario

The worker checks preference and a dedupe key before sending.

event(userId=42, type=order.shipped) -> queue
worker: 
  if channel disabled by preference: drop
  if dedupeKey seen: drop
  send via provider, on failure retry with backoff
Text

Walk through the example

The event is queued rather than sent inline, so the shipping service is not coupled to the email provider. The worker drops disabled channels and already-seen keys, then sends with retries. Each decision is a filter that reduces duplicate or unwanted messages.

Common mistake

Sending synchronously from the domain service, which couples latency and failures to the provider. Another is having no deduplication, so a reprocessed event spams the user, or ignoring preferences entirely.

Verify the behavior

Replay an event and confirm the dedupe key prevents a second send. Disable a channel and confirm no message goes out on it. Fail the primary provider and confirm the retry or fallback delivers, and that a digest batches multiple events into one message.

Interview exercise

Why queue notification events instead of sending from the service that raised them?

Answer and reasoning

Queuing decouples the producer from delivery, so a slow provider does not slow the domain request and a provider outage does not fail the business operation. It also enables retries, ordering and scaling independent of the producer, and it gives a durable record of what should be sent. The cost is eventual delivery, which is acceptable for notifications.

Continue learning

Compare fan-out in Feed fan-out and delivery in SQS delivery. Read the AWS messaging guidance and try the System design interview questions.

More in System Design

read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
read ✓System Design · hard

System Design: Idempotent APIs

Make retried requests safe with idempotency keys, store the result per key, and return the original response on a repeat.

~2 min readread →
esc