Ch. 19 · System Design

System Design Cache Stampedes and Refill Coordination

System Design Cache Stampedes and Refill Coordination. Learn the reasoning, a practical example, common mistakes and an interview exercise.

~2 min readbeginnerupdated Oct 3, 2026

A stampede occurs when many clients simultaneously refill the same missing or expired entry. Coordinate refill work and spread expiration.

Before you start

You should understand API requests, storage and basic capacity estimates. Begin with a concrete user action and its correctness requirement. Draw data flow and failure boundaries before selecting infrastructure; a technology name by itself does not explain why a design meets the requirement.

The practical goal is to reason through this situation: A popular key expires and hundreds of requests query the database together. Read the walkthrough first, then try the interview exercise before opening its answer. The important part is explaining the decision and its consequences, rather than remembering a definition alone.

Step-by-step walkthrough

Step 1: Identify synchronized misses

Many requests may arrive for one popular expired key.

Step 2: Coordinate refill work

Use a suitable single-flight or distributed coordination mechanism.

Step 3: Choose stale serving policy

Old data can protect the origin only where product correctness permits it.

Worked scenario

A popular key expires and hundreds of requests query the database together.

A popular product expires at noon and 500 requests all load it from the database. Increasing cache capacity does not prevent this shared-key miss. Coalescing refills reduces duplicate work; jitter spreads expirations across keys, while safe stale serving keeps users working during a bounded refill failure.

Common mistake

Increasing cache size does not prevent synchronized expiration.

Verify the behavior

Expire the key under concurrent load and count origin calls plus wait time.

Interview exercise

Protect the origin.

Answer and reasoning

Use single-flight or suitable locking, jittered expiry and a stale-serving policy where product correctness permits it.

Continue learning

Compare the scenario with the System Design interview questions and test your understanding with the System Design MCQs. For terminology and implementation details, consult the reference material.

More in System Design

read ✓System Design · easy

System Design Cache-Aside and Stale Data

System Design Cache-Aside and Stale Data. Learn the reasoning, a practical example, common mistakes and an interview exercise.

~2 min readread →
read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
esc