Backpressure is the mechanism that stops a fast producer from overwhelming a slow consumer. Without it, requests and messages accumulate in queues until memory is exhausted or latency collapses; with it, the system rejects or slows work before it degrades.
Before you start
You should understand queues, request lifecycles and basic capacity reasoning. This article is conceptual with small code sketches rather than a full implementation.
Step-by-step walkthrough
Step 1: Bound every queue
An unbounded queue hides overload until it fails catastrophically. Set a maximum size and a maximum wait time, and treat “queue full” as a designed response rather than an exception. A bound converts an outage into a controlled rejection.
Step 2: Signal the producer
When a consumer is saturated, tell the producer to slow down: return an HTTP 429 with Retry-After, use a reactive-streams request count, or apply TCP flow control. A signal that the producer ignores is not backpressure, so the protocol must be able to carry the signal.
Step 3: Shed load by priority
Under sustained pressure, reject the least important work first. Keep health checks and critical writes, and drop or defer background jobs and analytics. Load shedding is a policy decision made in advance, not an ad hoc error.
Worked scenario
A bounded worker returns a rejection instead of queueing without limit.
function createBoundedQueue(limit) {
const queue = [];
return {
push(task) {
if (queue.length >= limit) return { accepted: false, retryAfter: 1 };
queue.push(task);
return { accepted: true };
},
size: () => queue.length,
};
}
const q = createBoundedQueue(100);
console.log(q.push(() => {}).accepted); // trueWalk through the example
The queue caps at 100 and returns accepted: false with a retry hint once full. The caller can surface a 429 and a Retry-After header, so the client slows down. Nothing grows without bound, and the rejection path is explicit rather than an out-of-memory crash later.
Common mistake
Using an unbounded in-memory queue and assuming the broker or database will absorb the spike; it will not, and latency grows unbounded first. Another is returning 500 for overload instead of 429, which tells clients to retry immediately and deepens the overload.
Verify the behavior
Load-test past the queue bound and assert rejections are fast and well-formed rather than slow failures. Confirm Retry-After is present and that a well-behaved client backs off. Check that memory stays flat under sustained overload and that prioritized endpoints keep succeeding.
Interview exercise
A checkout service and an analytics pipeline share one worker pool. How do you keep checkout responsive during a traffic spike?
Answer and reasoning
Separate the pools or give checkout a reserved capacity and priority, then bound the analytics queue so it sheds first when saturated. Add admission control that rejects excess analytics work early with a 429. This isolates the critical path so background load cannot starve checkout, which is the core of bulkhead-style isolation plus backpressure.
Continue learning
Compare isolation patterns in Bulkhead isolation and rate limiting in Rate limits. Read the AWS backpressure and load shedding guidance and try the System design interview questions.