“How would you add a Redis cache in front of this database?” is one of the most frequent backend and system design prompts. A good answer names a pattern (usually cache-aside), explains how writes keep the cache honest, and then handles the question the interviewer is waiting for: “what happens when a popular key expires under heavy traffic?” That last part, the cache stampede or thundering herd, is where most answers fall short.
This guide walks through the caching patterns with Redis 7.4/8.x and Python (redis-py), the correct invalidation order, and tested code for stampede protection. The Python examples were run with redis-py against fakeredis, an in-process Redis emulator.
Before you start
You should know basic GET, SET ... EX and DEL, and what a TTL is. If expiry and eviction are new, read Redis eviction policies and key expiry first. It also helps to have a concrete backing store in mind, such as PostgreSQL, where a query costs tens or hundreds of milliseconds and the database can handle far fewer concurrent queries than Redis.
The short answer
I would use cache-aside: the application reads from Redis, and on a miss loads from the database and writes the value back with a TTL. On writes I update the database, then delete the cache key, so the next read loads the committed value; the TTL bounds any staleness left by races. Write-through (write cache and database together) keeps the cache warm at the cost of slower writes, and write-behind (write to Redis, flush to the database later) is fast but can lose data. For hot keys I prevent a stampede with a short rebuild lock (SET NX PX) so one request reloads while others wait or serve stale data, plus TTL jitter and optionally early probabilistic refresh.
How it works
| Pattern | Read path | Write path | Strength | Weakness |
|---|---|---|---|---|
| Cache-aside | app: cache, then DB on miss, then populate | app: update DB, delete key | simple, survives cache outage | miss penalty, small stale window |
| Read-through | cache layer loads misses itself | usually combined with another pattern | app code stays simple | needs a loader-aware cache library |
| Write-through | from cache | write cache and DB synchronously | cache always warm for written keys | slower writes, caches unread data |
| Write-behind | from cache | write cache, flush DB asynchronously | very fast writes, batched DB load | data loss if Redis fails before the flush |
Why delete instead of update on write
Suppose two requests update the same product price. Writer A commits 39 and writer B commits 41, so the database holds 41. If each writer then sets the cache, network timing can make A’s SET arrive after B’s, and Redis holds 39 until the TTL expires. A DEL cannot be reordered into a wrong value: whichever arrives last, the key is gone, and the next reader loads 41.
The order matters too. Delete, then update the database leaves a gap in which a reader misses, loads the old row and repopulates the cache just before the database changes. Update the database, then delete is much safer. A narrow race remains: a reader that loaded the old row before the commit can write it back after the delete. Short TTLs bound that, a second delayed delete (for example 500 ms later) shrinks it, and for strict needs you drive invalidation from the database’s change log (CDC) instead of application code.
Why stampedes happen
A key read 5,000 times a second with a 60-second TTL is rebuilt once a minute. At the moment it expires, every in-flight request misses at once. If a rebuild takes 300 ms, around 1,500 requests start the same expensive query before the first one finishes. The database slows down, rebuilds take longer, and more requests pile up.
Step-by-step walkthrough
Step 1: Implement cache-aside with a TTL
import json
import redis
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
def get_product(pid: int, ttl: int = 300) -> dict:
key = f"product:{pid}"
cached = r.get(key)
if cached is not None:
return json.loads(cached)
product = load_product_from_db(pid) # SELECT ... WHERE id = %s
r.set(key, json.dumps(product), ex=ttl)
return product
def update_price(pid: int, price: int) -> None:
save_price_to_db(pid, price) # commit first
r.delete(f"product:{pid}") # then invalidateIf Redis is down, wrap the cache calls in try/except and fall back to the database, so the cache stays an optimization rather than a dependency that can take reads offline.
Step 2: Cache negative results and add jitter
Requests for IDs that do not exist miss every time and reach the database (cache penetration). Cache a short-lived sentinel, and randomize TTLs so keys loaded together do not expire together:
import random
NONE = "__none__"
def cache_set(key: str, value: str | None, ttl: int) -> None:
if value is None:
r.set(key, NONE, ex=30) # short negative TTL
else:
r.set(key, value, ex=ttl + random.randint(0, ttl // 5))The read path treats NONE as “not found” without querying the database.
Step 3: Coalesce rebuilds with a lock
Only one request should rebuild a missing hot key. The others wait briefly and re-check the cache. The lock is a key with a random token and an expiry, released by a Lua compare-and-delete so a slow holder never deletes someone else’s lock:
import time, uuid
RELEASE = """
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0
"""
def get_product(pid: int, ttl: int = 300) -> dict:
key = f"product:{pid}"
cached = r.get(key)
if cached is not None:
return json.loads(cached)
lock_key, token = f"lock:{key}", str(uuid.uuid4())
for _ in range(50): # wait up to about 2.5 s
if r.set(lock_key, token, nx=True, px=5000):
try:
cached = r.get(key) # filled while we waited?
if cached is not None:
return json.loads(cached)
product = load_product_from_db(pid)
r.set(key, json.dumps(product), ex=ttl + random.randint(0, ttl // 5))
return product
finally:
r.eval(RELEASE, 1, lock_key, token)
time.sleep(0.05)
cached = r.get(key)
if cached is not None:
return json.loads(cached)
return load_product_from_db(pid) # degrade rather than failIn a test with 100 threads requesting the same missing key and a loader that sleeps 200 ms, the naive version called the loader 100 times; this version called it once.
Step 4: Refresh early instead of waiting for expiry
Locks still make some requests wait. Probabilistic early expiration (the XFetch algorithm) lets each reader decide to refresh slightly before the deadline, with a probability that rises as the deadline approaches and with how long the rebuild takes. Store the value, the rebuild time and a logical expiry in a hash:
import math
def get_with_early_refresh(key: str, loader, ttl: int = 300, beta: float = 1.0):
data = r.hgetall(key)
now = time.time()
if data:
expiry, delta = float(data["expiry"]), float(data["delta"])
# -log(random()) is positive; slow rebuilds start refreshing earlier
if now - delta * beta * math.log(random.random()) < expiry:
return json.loads(data["value"])
start = time.time()
value = loader()
delta = time.time() - start
pipe = r.pipeline()
pipe.hset(key, mapping={"value": json.dumps(value), "delta": delta, "expiry": time.time() + ttl})
pipe.expire(key, ttl)
pipe.execute()
return valueUsually one request refreshes a little early while everyone else keeps getting cached data, so the key never actually disappears under load. The same structure supports stale-while-revalidate: give the Redis key a longer real TTL than the logical expiry, serve the stale value after the logical deadline, and refresh in the background.
Step 5: Choose write-through or write-behind only when needed
Write-through suits data that is read immediately after it is written, such as a user’s own profile. Write-behind suits high-volume counters where a little loss is tolerable. Make the buffer durable by writing changes to a Stream and flushing it with a consumer group:
r.hincrby(f"likes:{post_id}", "count", 1) # fast path for readers
r.xadd("likes:changes", {"post": post_id, "delta": 1}, maxlen=1_000_000, approximate=True)
# a worker reads with XREADGROUP, aggregates per post, writes to the DB, then XACKsWorked scenario
A retail homepage shows “featured products”, built by a 400 ms query that joins inventory, pricing and promotions. It is cached under one key with a fixed 60-second TTL and read about 3,000 times a second at peak. Every minute, database CPU spiked to 100% for several seconds, the homepage p99 jumped to 8 seconds, and the connection pool emptied.
# Broken: every concurrent miss runs the 400 ms query
def featured():
cached = r.get("home:featured")
if cached:
return json.loads(cached)
rows = run_featured_query()
r.set("home:featured", json.dumps(rows), ex=60)
return rowsAt 3,000 requests a second, roughly 1,200 requests arrive during the 400 ms rebuild, and each one starts its own query. The fix combined two techniques: the lock-based rebuild from Step 3, so only one query runs, and a background refresher that rewrites the key every 30 seconds while it has a 120-second TTL, so in normal operation it never expires at all. The lock path is the safety net after a Redis restart or a failover. Database CPU on the minute mark dropped to its baseline, and the homepage p99 stayed under 100 ms.
Common mistake
- “Update the cache on every write.” Concurrent writers can leave the older value cached. Delete instead.
- “Delete the key, then update the database.” A reader can repopulate the old value in between.
- “TTL solves consistency.” It bounds staleness; it does not prevent it.
- Locks without expiry or token. A crashed rebuilder blocks the key forever, and a bare
DELcan remove another request’s lock. - Same TTL for everything loaded at once. Mass expiry becomes an avalanche; add jitter.
- Caching unbounded result sets. A list of “all products” becomes a big key that is expensive to rebuild and to invalidate.
Verify the behavior
Count loader calls under concurrency, which is how the numbers above were produced:
import threading
calls = 0
def load_product_from_db(pid):
global calls
calls += 1
time.sleep(0.2)
return {"id": pid}
r.delete("product:42")
threads = [threading.Thread(target=get_product, args=(42,)) for _ in range(100)]
for t in threads: t.start()
for t in threads: t.join()
print(calls) # 1 with the lock; 100 with the naive versionIn production, watch keyspace_hits and keyspace_misses from INFO stats, database queries per second for the cached query, and latency percentiles around TTL boundaries. A sawtooth in database load aligned with the TTL is the signature of a stampede.
Follow-up questions
What happens if the lock holder crashes mid-rebuild? The lock expires after its PX timeout and the next request takes over; waiters fall back to the database after their own timeout rather than hanging.
How do you invalidate a cached list when one item changes? Either delete the list key too, or cache lists as IDs only and fetch the items by key, so changing an item never stales the list.
How does client-side caching change this? Redis 6+ supports CLIENT TRACKING: the server pushes invalidation messages when keys a client has read change, so an in-process cache can be kept consistent with Redis.
When is caching the wrong answer? When data must be read-your-writes consistent across users, when the hit ratio would be low, or when the query can be made cheap with an index.
Interview exercise
An inventory service caches stock levels with cache-aside: on purchase it decrements stock in PostgreSQL and then runs SET stock:{sku} <new value> EX 600. During a flash sale, the site keeps showing items as in stock after they sold out, sometimes for minutes. Explain the bug and redesign the flow.
Answer and reasoning
Concurrent purchases commit in one order but their SETs can reach Redis in another, so an older, higher stock value can overwrite a newer one and live for the 10-minute TTL. Switch the write path to update-then-DEL, so the next read loads the committed value. During a flash sale the key is hot, so pair the delete with a rebuild lock to avoid a stampede on every purchase, or shorten the TTL to a few seconds. If Redis is meant to be the fast path for stock, make it authoritative for the sale instead: keep the counter in Redis, decrement it atomically with a Lua script that refuses to go below zero, and write the result to the database asynchronously. Showing slightly stale availability is acceptable; overselling is not, so the final check must happen in the store that is authoritative.
Continue learning
Practise with the Redis interview questions and the Redis MCQs. Related notes: System design cache-aside and stale data, System design cache stampedes and refill coordination and Node.js caching layers and invalidation. Primary sources: the Redis documentation, the SET command reference for NX, PX and KEEPTTL, and the client-side caching reference.