Ch. 19 · System Design

System Design: Backpressure and Load Shedding

Protect a service from overload by bounding queues, signaling producers to slow down, and shedding low-priority work.

~2 min readadvancedupdated Oct 5, 2026

Backpressure is the mechanism that stops a fast producer from overwhelming a slow consumer. Without it, requests and messages accumulate in queues until memory is exhausted or latency collapses; with it, the system rejects or slows work before it degrades.

Before you start

You should understand queues, request lifecycles and basic capacity reasoning. This article is conceptual with small code sketches rather than a full implementation.

Step-by-step walkthrough

Step 1: Bound every queue

An unbounded queue hides overload until it fails catastrophically. Set a maximum size and a maximum wait time, and treat “queue full” as a designed response rather than an exception. A bound converts an outage into a controlled rejection.

Step 2: Signal the producer

When a consumer is saturated, tell the producer to slow down: return an HTTP 429 with Retry-After, use a reactive-streams request count, or apply TCP flow control. A signal that the producer ignores is not backpressure, so the protocol must be able to carry the signal.

Step 3: Shed load by priority

Under sustained pressure, reject the least important work first. Keep health checks and critical writes, and drop or defer background jobs and analytics. Load shedding is a policy decision made in advance, not an ad hoc error.

Worked scenario

A bounded worker returns a rejection instead of queueing without limit.

function createBoundedQueue(limit) {
  const queue = [];
  return {
    push(task) {
      if (queue.length >= limit) return { accepted: false, retryAfter: 1 };
      queue.push(task);
      return { accepted: true };
    },
    size: () => queue.length,
  };
}
const q = createBoundedQueue(100);
console.log(q.push(() => {}).accepted); // true
JavaScript

Walk through the example

The queue caps at 100 and returns accepted: false with a retry hint once full. The caller can surface a 429 and a Retry-After header, so the client slows down. Nothing grows without bound, and the rejection path is explicit rather than an out-of-memory crash later.

Common mistake

Using an unbounded in-memory queue and assuming the broker or database will absorb the spike; it will not, and latency grows unbounded first. Another is returning 500 for overload instead of 429, which tells clients to retry immediately and deepens the overload.

Verify the behavior

Load-test past the queue bound and assert rejections are fast and well-formed rather than slow failures. Confirm Retry-After is present and that a well-behaved client backs off. Check that memory stays flat under sustained overload and that prioritized endpoints keep succeeding.

Interview exercise

A checkout service and an analytics pipeline share one worker pool. How do you keep checkout responsive during a traffic spike?

Answer and reasoning

Separate the pools or give checkout a reserved capacity and priority, then bound the analytics queue so it sheds first when saturated. Add admission control that rejects excess analytics work early with a 429. This isolates the critical path so background load cannot starve checkout, which is the core of bulkhead-style isolation plus backpressure.

Continue learning

Compare isolation patterns in Bulkhead isolation and rate limiting in Rate limits. Read the AWS backpressure and load shedding guidance and try the System design interview questions.

More in System Design

read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
read ✓System Design · hard

System Design: Idempotent APIs

Make retried requests safe with idempotency keys, store the result per key, and return the original response on a repeat.

~2 min readread →
esc