Senior Redis interviews almost always reach scaling and availability: “how does Redis replication work?”, “Sentinel or Cluster, and why?”, “what are hash slots?”, “what is the difference between MOVED and ASK?”, and the sharp one, “can Redis lose a write it acknowledged?” These questions test whether you understand Redis’s consistency model, not just its commands.
This guide covers replication, Sentinel and Cluster in Redis 7.4 and 8.x, with the commands to observe each and a failure scenario that shows why the answer to the last question is yes. Valkey implements the same protocols.
Before you start
You should be comfortable with basic Redis commands and know what a primary/replica setup is in any database. Two distributed systems ideas help: a network partition (some nodes can reach each other but not the rest) and a quorum (a minimum number of voters that must agree). If persistence is unfamiliar, read Redis persistence: RDB vs AOF first, because full resynchronization uses RDB snapshots.
The short answer
Redis replication is asynchronous: a replica connects to the primary, continues from its replication ID and offset (PSYNC), and receives the stream of writes, but the primary acknowledges clients without waiting for replicas. Sentinel adds automatic failover for one primary and its replicas: Sentinels agree the primary is down, elect a leader by majority, and promote a replica. Redis Cluster shards the keyspace into 16,384 hash slots (CRC16(key) mod 16384) spread over several primaries, each with replicas; clients route commands with a slot map and follow MOVED (slot moved permanently) or ASK (slot mid-migration) redirects. In all three, a failover can lose recently acknowledged writes.
How it works
Replication
A replica runs REPLICAOF <host> <port> and sends PSYNC <replid> <offset>. The primary keeps a circular replication backlog (repl-backlog-size, 1 MB by default). If the requested offset is still in the backlog, the primary sends just the missing bytes (partial resync). Otherwise it performs a full resync: it produces an RDB snapshot (streamed directly over the socket by default since Redis 7.0, repl-diskless-sync yes), sends it, and then streams the writes buffered meanwhile. After that, every write command the primary executes is forwarded to replicas.
Because the client’s OK is sent before replicas receive the write, a primary that crashes right after replying can take that write with it. WAIT 1 100 blocks until at least one replica acknowledges the write or 100 ms pass, and WAITAOF (7.2) also waits for fsync; both reduce the window but do not make failover safe in every case, because the promoted replica might not be the one that acknowledged.
Sentinel
Each Sentinel pings the primary and replicas. No valid reply for down-after-milliseconds makes that Sentinel consider the primary subjectively down (SDOWN). When at least quorum Sentinels report SDOWN, it becomes objectively down (ODOWN). To act, one Sentinel must be elected leader by a majority of all Sentinels; it chooses a replica (by replica-priority, then replication offset, then run ID), promotes it with REPLICAOF NO ONE, repoints the other replicas and announces the new address. Clients ask Sentinels for the current primary rather than hard-coding it.
Cluster
| Concept | Detail |
|---|---|
| Slot of a key | CRC16(key) mod 16384; CLUSTER KEYSLOT foo returns 12182 |
| Hash tag | only the text inside the first {...} is hashed: {user:42}:cart |
| Routing | clients cache the slot map (CLUSTER SHARDS) and connect to each primary directly |
MOVED slot host:port |
slot belongs elsewhere now; retry there and refresh the map |
ASK slot host:port |
slot is migrating and this key already moved; send ASKING then the command, once |
| Failure detection | gossip over the cluster bus (data port + 10000); replica promoted when a majority of primaries agree |
| Limits | database 0 only; multi-key operations need one slot (CROSSSLOT otherwise) |
16,384 slots is a deliberate size: enough granularity to balance across up to around 1,000 primaries, while the slot bitmap in each gossip heartbeat stays at 2 KB.
Step-by-step walkthrough
Step 1: Observe replication
# on the replica
redis-cli -p 6380 REPLICAOF 127.0.0.1 6379
# on the primary
redis-cli -p 6379 INFO replication
# role:master
# connected_slaves:1
# slave0:ip=127.0.0.1,port=6380,state=online,offset=4821,lag=0
# master_replid:8f1c...
# master_repl_offset:4821
# repl_backlog_size:1048576INFO still uses the historic master/slave field names for compatibility. Compare master_repl_offset with each replica’s offset to measure lag in bytes. On a busy primary, raise repl-backlog-size so a replica that disconnects for a few seconds can partially resync instead of forcing a full snapshot.
Step 2: Add Sentinel and connect through it
# sentinel.conf, on three separate hosts
sentinel monitor mymaster 10.0.0.1 6379 2
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 60000
sentinel parallel-syncs mymaster 1from redis.sentinel import Sentinel
sentinel = Sentinel([("10.0.1.1", 26379), ("10.0.1.2", 26379), ("10.0.1.3", 26379)],
socket_timeout=0.5)
primary = sentinel.master_for("mymaster", socket_timeout=0.5)
replica = sentinel.slave_for("mymaster", socket_timeout=0.5) # reads may be stale
primary.set("greeting", "hello")With quorum 2 of 3, two Sentinels must agree the primary is down, and two (a majority) must elect the leader. With only two Sentinels, losing one host would make failover impossible.
Step 3: Create a cluster and watch keys spread
redis-cli --cluster create 10.0.0.1:7000 10.0.0.2:7000 10.0.0.3:7000 \
10.0.0.4:7000 10.0.0.5:7000 10.0.0.6:7000 --cluster-replicas 1
redis-cli -c -h 10.0.0.1 -p 7000
10.0.0.1:7000> CLUSTER KEYSLOT foo
(integer) 12182
10.0.0.1:7000> SET foo bar
-> Redirected to slot [12182] located at 10.0.0.3:7000
OKThree primaries each own roughly 5,461 slots, and each has one replica. -c makes redis-cli follow redirects; without it you would see the raw MOVED error. Application clients do this transparently:
from redis.cluster import RedisCluster
rc = RedisCluster(host="10.0.0.1", port=7000)
rc.set("foo", "bar") # sent straight to the owner of slot 12182Step 4: Group related keys with hash tags
10.0.0.1:7000> MSET user:42:name Ana user:42:plan pro
(error) CROSSSLOT Keys in request don't hash to the same slot
10.0.0.1:7000> MSET {user:42}:name Ana {user:42}:plan pro
OKBoth tagged keys hash only user:42, so they share a slot and can be used together in MULTI, Lua scripts and multi-key commands. Tag by the entity that needs atomic updates; a single global tag would put all data on one shard.
Step 5: Choose the topology
Use a single primary with replicas and Sentinel when the dataset and write rate fit on one machine, and you value multi-key freedom (any keys in one transaction or script). Use Cluster when memory or write throughput exceeds one primary, accepting the same-slot rule and a cluster-aware client. Both are managed for you by most cloud offerings, but the consistency model is the same.
Worked scenario
A team stored payment idempotency keys in Redis with Sentinel: one primary, two replicas, three Sentinels. A switch failure partitioned the primary and two application servers from everything else. On the majority side, Sentinels marked the primary ODOWN, elected a leader and promoted a replica within about 10 seconds. On the minority side, the old primary was still running, and the two application servers kept writing to it for six minutes. When the network healed, Sentinel reconfigured the old primary as a replica of the new one, and it discarded its dataset to resync. Six minutes of idempotency keys vanished, and some retried payments were processed twice.
# Before: the isolated primary accepts writes indefinitely
min-replicas-to-write 0The fix makes an isolated primary stop accepting writes once it cannot see enough replicas:
# After: refuse writes unless at least 1 replica acknowledged within 10 seconds
min-replicas-to-write 1
min-replicas-max-lag 10Now the isolated primary rejects writes with a NOREPLICAS error after about 10 seconds, bounding the loss to that window. The team also moved the final duplicate check to the payments database, because Redis with asynchronous replication cannot be the only guard for money.
Common mistake
- “Replicas make Redis strongly consistent.” Replication is asynchronous;
WAITnarrows the gap but does not give linearizability. - “Two Sentinels are enough.” Failover needs a majority; with two, losing one blocks failover. Use at least three, on separate failure domains.
- “Cluster proxies requests to the right node.” It redirects; the client must follow
MOVEDandASK. - Confusing
MOVEDandASK. Updating the slot map onASKsends future traffic to a node that does not own the slot yet. - One hash tag for everything. All keys land in one slot, so one shard does all the work.
- Using
SELECT 1in Cluster. Only database 0 exists; use key prefixes instead.
Verify the behavior
Make failover visible on a test setup:
redis-cli -p 26379 SENTINEL get-master-addr-by-name mymaster
# 1) "10.0.0.1"
# 2) "6379"
redis-cli -p 6379 DEBUG SLEEP 30 # primary stops answering (needs enable-debug-command)
redis-cli -p 26379 SENTINEL get-master-addr-by-name mymaster
# after down-after-milliseconds plus election: a replica's address
# Cluster: confirm slot coverage and simulate a failover
redis-cli --cluster check 10.0.0.1:7000
redis-cli -h 10.0.0.4 -p 7000 CLUSTER FAILOVER # run on a replica: graceful promotion
redis-cli -h 10.0.0.1 -p 7000 CLUSTER SHARDSTo prove asynchronous loss, write a counter in a loop, kill the primary with kill -9, and compare the last acknowledged value with the value on the promoted replica: the replica is often a few increments behind.
Follow-up questions
What does cluster-require-full-coverage control? With yes (default), the whole cluster stops serving if any slot has no owner; with no, nodes keep serving the slots they still cover.
How do you read from replicas in Cluster? Connect with READONLY on a replica connection (or enable replica reads in the client) and accept stale reads.
What happens to Pub/Sub in Cluster? Classic PUBLISH is broadcast to every node; Redis 7 sharded Pub/Sub (SSUBSCRIBE, SPUBLISH) keeps a channel on the shard that owns its slot.
How does resharding stay online? Slots migrate key by key with MIGRATE; during that time the source answers ASK for keys already moved, and ownership flips when the slot is empty.
Interview exercise
An e-commerce platform runs one Redis primary (48 GB, near its memory limit) with two replicas and Sentinel. Peak writes are 180,000 per second and the primary’s CPU is near 100%. The application uses Lua scripts that touch a user’s cart, wishlist and a global stats:orders counter in one call. The team wants to move to Redis Cluster. What must change?
Answer and reasoning
Cluster is the right direction, since both memory and write throughput exceed one primary, but the scripts break: a script’s keys must all be in one slot, and the global counter will live on a different shard from most users. Re-key per-user data with a hash tag ({user:42}:cart, {user:42}:wishlist) so each user’s keys share a slot, and move the global counter out of the script: increment it in a separate command, or shard it as stats:orders:{0..15} with a random suffix and sum the shards on read, which also avoids a hot key. Every script must declare its keys in KEYS so the client can route it. Switch to a cluster-aware client, remove any SELECT usage, and test behavior during resharding (ASK, TRYAGAIN). Replication stays asynchronous, so the durability expectations do not change.
Continue learning
Practise with the Redis interview questions and the Redis MCQs. Related notes: Redis persistence: RDB vs AOF and Redis distributed locks and Redlock, where asynchronous failover is the key weakness. Primary sources: Redis replication, High availability with Redis Sentinel, Scale with Redis Cluster and the Redis cluster specification.