Retries turn a transient failure such as a dropped connection into a success, but a naive retry loop can make an outage worse. Only retry operations that are safe to repeat, wait longer after each failure, add jitter so retries do not synchronize, and cap both attempts and total time.
Before you start
You should be comfortable with async/await and error handling. This article covers client-side retries; it assumes you can tell which errors are transient.
Step-by-step walkthrough
Step 1: Retry only transient, idempotent operations
Retry timeouts, connection resets and 503/429 responses. Do not retry a 400, a validation error, or a non-idempotent write without an idempotency key, or you risk duplicating the side effect. Classify the error before deciding.
Step 2: Back off with jitter
Wait longer after each attempt, for example 2 ** attempt * base, and add a random fraction so a fleet of clients does not retry at the same instant. Without jitter, synchronized retries create a thundering herd that re-breaks the recovering service.
Step 3: Cap attempts and total time
Set a maximum attempt count and a deadline, and honor a server Retry-After when present. An uncapped loop can pile up work and outlive the request that needed it, so the budget is part of the contract.
Worked scenario
The loop backs off with cap and jitter and rethrows the last error.
async function retry(fn, attempts = 3) {
let lastError;
for (let i = 0; i < attempts; i += 1) {
try {
return await fn();
} catch (error) {
lastError = error;
const delay = Math.min(2 ** i * 100 + Math.random() * 100, 2000);
await new Promise((resolve) => setTimeout(resolve, delay));
}
}
throw lastError;
}Walk through the example
Each failure waits exponentially longer, bounded at two seconds, with up to a hundred milliseconds of jitter to desynchronize callers. After the attempt budget is exhausted, the last error is thrown so the caller still sees the failure. The function retries blindly, so use it only with idempotent operations.
Common mistake
Retrying every error, including deterministic ones, which wastes time and can duplicate writes. Another is exponential backoff without jitter across many clients, which concentrates retries and overloads the recovering dependency.
Verify the behavior
Assert that a transient failure succeeds within the budget and a permanent failure stops immediately. Measure the delays and confirm they grow and include jitter. Confirm a non-idempotent call is guarded by an idempotency key before retrying it.
Interview exercise
Why is jitter necessary when many clients retry the same service?
Answer and reasoning
Without jitter, every client that failed at the same time waits the same interval and retries simultaneously, hitting the recovering service in a synchronized spike that can cause another failure. Jitter spreads attempts across the interval, smoothing the load so the service can recover. Backoff controls the rate; jitter controls the correlation.
Continue learning
Compare resilience patterns in Timeout phases and Retry budgets. Read the AWS exponential backoff and jitter guidance and try the Node.js interview questions.