“How would you implement a distributed lock with Redis?” is a favourite because it starts simple and gets deep fast. The first answer (SETNX) is usually wrong in a small way; the follow-ups (“what if the holder crashes?”, “what if the work takes longer than the TTL?”, “what is Redlock and is it safe?”) test whether you understand that a lease-based lock in a distributed system can never be perfectly exclusive on its own.
This guide builds a single-instance lock step by step for Redis 7.4/8.x, adds renewal and fencing tokens, and then covers Redlock and the debate around it. The Python lock class and the fencing check were run with redis-py against fakeredis and SQLite.
Before you start
You should know SET with its options, what a TTL is, and that Redis executes one command (or one Lua script) at a time. It helps to know what a lease is: a lock that automatically expires, so a crashed holder cannot block others forever. Familiarity with process pauses (garbage collection, a VM being suspended) matters for the last sections. For locking in general, see Microservices distributed locking.
The short answer
I acquire with one atomic command, SET lock:resource <random-token> NX PX 30000: NX makes only one client succeed, PX makes the lock expire if the holder dies, and the token identifies the owner. I release with a Lua script that deletes the key only if it still holds my token, so I never delete a lock someone else acquired after mine expired. If work can outlast the TTL, I renew it while I still own it. Because a paused client can still act after its lock expired, anything that must be correct (not just efficient) also needs a fencing token checked by the resource. Redlock extends this to several independent Redis primaries, but Martin Kleppmann’s critique shows it still depends on timing assumptions.
How it works
The lock is just a key. Its existence means “held”, its value says by whom, and its TTL bounds how long a dead holder can block others.
127.0.0.1:6379> SET lock:{report:daily} 7f9c2b1e NX PX 30000
OK
127.0.0.1:6379> SET lock:{report:daily} a41d77c0 NX PX 30000
(nil)
127.0.0.1:6379> PTTL lock:{report:daily}
(integer) 29412Why each part matters:
| Part | Without it |
|---|---|
One SET ... NX PX command |
SETNX then EXPIRE can crash in between, leaving a lock that never expires |
PX expiry |
a crashed holder blocks the resource forever |
| Random token per acquisition | the holder cannot tell whether the lock is still its own |
| Token-checked release (Lua) | a slow holder’s DEL removes the next holder’s lock |
| Renewal or a generous TTL | long work silently loses the lock |
| Fencing token | a paused holder can still write after losing the lock |
The release must compare and delete atomically. A GET then DEL from the client can interleave: your GET sees your token, the lock expires, another client acquires it, and your DEL removes theirs. A Lua script runs as one unit, so nothing can happen between the comparison and the delete.
The braces in lock:{report:daily} are a Cluster hash tag. They are not needed on a single instance, but they keep the lock key and its fencing counter in the same slot, so one script can touch both.
Step-by-step walkthrough
Step 1: Acquire atomically and get a fencing token
The acquire script sets the lock and, only on success, increments a counter that produces a strictly increasing fencing token:
-- KEYS[1] = lock key, KEYS[2] = fence counter, ARGV[1] = token, ARGV[2] = ttl ms
if redis.call('SET', KEYS[1], ARGV[1], 'NX', 'PX', ARGV[2]) then
return redis.call('INCR', KEYS[2])
end
return 0SET ... NX returns a status reply on success, which Lua sees as a table (truthy), and a nil reply on failure, which Lua sees as false.
Step 2: Release and renew only while you own the lock
-- release: KEYS[1] = lock key, ARGV[1] = token
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0-- extend: KEYS[1] = lock key, ARGV[1] = token, ARGV[2] = ttl ms
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('PEXPIRE', KEYS[1], ARGV[2])
end
return 0A renewal that returns 0 means the lock is gone: stop work rather than carry on. Redisson’s watchdog uses the same idea, renewing a 30-second lease every 10 seconds while the holding thread is alive.
Step 3: Wrap it in a small class
import uuid
import redis
class RedisLock:
def __init__(self, r: redis.Redis, name: str, ttl_ms: int = 30_000):
self.r, self.ttl_ms = r, ttl_ms
self.key, self.fence_key = f"lock:{{{name}}}", f"fence:{{{name}}}"
self.token = str(uuid.uuid4())
self._acquire = r.register_script(ACQUIRE)
self._extend = r.register_script(EXTEND)
self._release = r.register_script(RELEASE)
self.fence = None
def acquire(self) -> bool:
fence = self._acquire(keys=[self.key, self.fence_key], args=[self.token, self.ttl_ms])
if fence:
self.fence = fence
return True
return False
def extend(self) -> bool:
return self._extend(keys=[self.key], args=[self.token, self.ttl_ms]) == 1
def release(self) -> bool:
return self._release(keys=[self.key], args=[self.token]) == 1register_script sends EVALSHA and falls back to EVAL if the server has not cached the script yet. In a test with a 200 ms TTL, client A acquired (fence 1), slept 300 ms, and client B then acquired (fence 2); A’s extend() and release() both returned False, and B’s lock was untouched.
Step 4: Make the resource reject stale holders
The lock alone cannot stop a client that was paused past its TTL. The fencing token can, if the storage checks it:
def save_report(db, body: str, fence: int) -> bool:
cur = db.execute(
"UPDATE reports SET body = ?, fence = ? WHERE name = 'daily' AND fence < ?",
(body, fence, fence),
)
db.commit()
return cur.rowcount == 1 # False: a newer lock holder already wroteWhen B (fence 2) writes first and A (fence 1) wakes up and writes, A’s update matches no rows. This is the property Kleppmann argues a lock service must provide for correctness.
Worked scenario
A nightly job builds a large CSV and emails it to finance. Several instances run the scheduler, so a Redis lock ensures only one runs it:
# Broken
if r.set("lock:report", "1", nx=True, ex=10):
build_and_email_report() # usually 6 s, sometimes 40 s
r.delete("lock:report")One night the database was slow and the job took 40 seconds. At 10 seconds the lock expired, a second instance acquired it and started the same job, and at 40 seconds the first instance’s DEL removed the second instance’s lock, letting a third start. Finance received three emails, and two instances wrote the same output file concurrently, corrupting it.
The fix used every layer from the walkthrough:
lock = RedisLock(r, "report:daily", ttl_ms=30_000)
if lock.acquire():
try:
for chunk in build_report_chunks():
if not lock.extend(): # renew between chunks; stop if lost
raise RuntimeError("lock lost, aborting")
write_chunk(chunk)
save_report(db, final_path, lock.fence) # fenced write
send_email_once(report_date) # idempotent: keyed by date
finally:
lock.release()The token-checked release stopped instances deleting each other’s locks, renewal kept the lock during slow runs, the fenced write rejected any stale holder, and an idempotency key on the email made the final side effect safe even if everything else failed.
Common mistake
SETNXfollowed byEXPIRE. Two commands; a crash between them creates a lock with no expiry.- Releasing with
DEL. Deletes whoever holds the lock now, not necessarily you. - A fixed value like
"1". Without a unique token you cannot check ownership. - “The TTL guarantees exclusivity.” A GC pause, a slow network or a suspended VM can outlast it.
- “Replicas make the lock safe.” With asynchronous replication, a failover can promote a replica that never saw the lock key, so two clients hold it.
- Using locks where atomic commands suffice. Many “lock then update” flows are a single
INCR,HINCRBYor Lua script.
Verify the behavior
Reproduce the expiry race deliberately with a short TTL:
a = RedisLock(r, "report:daily", ttl_ms=200)
b = RedisLock(r, "report:daily", ttl_ms=200)
print(a.acquire(), a.fence) # True 1
print(b.acquire()) # False
time.sleep(0.3) # A "pauses" past its TTL
print(b.acquire(), b.fence) # True 2
print(a.extend(), a.release()) # False False: A cannot touch B's lock
print(r.get(b.key) == b.token) # TrueThese outputs match the run described above. In redis-cli, MONITOR on a test instance shows each EVALSHA and SET ... NX PX, and PTTL on the lock key confirms renewals.
Follow-up questions
What is Redlock? With N independent primaries (typically five, no replication between them), a client records the start time, tries SET NX PX on each with a short timeout, and holds the lock only if a majority (3 of 5) succeeded and the elapsed time is less than the TTL; the remaining validity is the TTL minus elapsed time and a drift allowance. On failure it releases on every node.
What is Kleppmann’s critique? A lease cannot guarantee mutual exclusion when clients can pause past the TTL, and Redlock relies on bounded clock drift and network delay. It also generates no fencing token. For efficiency locks a single Redis instance is enough; for correctness use fencing with a consensus system such as ZooKeeper or etcd. Antirez replied that Redlock’s assumptions are reasonable in practice, and the disagreement is about which guarantees you need.
How do you pick the TTL? Longer than the normal work plus margin, short enough that a crashed holder blocks others only briefly, with renewal for long work.
When would you not use Redis for a lock? When correctness depends on it and the resource cannot check fencing tokens; a database row lock (SELECT ... FOR UPDATE) inside the same transaction as the write is simpler and safer.
Interview exercise
A ticketing service holds a seat with SET seat:{event}:{seat} <userId> NX PX 600000 while the user pays, and confirms the booking in PostgreSQL after payment. The Redis primary fails over during a busy sale, and some seats are sold twice. Explain how, and fix the design.
Answer and reasoning
The seat holds were written to the old primary and acknowledged, but replication is asynchronous; the promoted replica had not received some of them, so other users could acquire the same seats with SET NX and both later paid. Redis is acting as the only guard for a correctness property, which it cannot guarantee across failover. Keep Redis for the fast, user-facing hold, but make PostgreSQL authoritative: a unique constraint on (event_id, seat_id) in the bookings table, or a conditional update that only succeeds while the seat is free, so the second confirmation fails cleanly and that user gets a refund or a new seat. Optionally use WAIT 1 50 after taking a hold to reduce, not eliminate, the window, and a fencing token derived from the booking version if holds are checked elsewhere.
Continue learning
Practise with the Redis interview questions and the Redis MCQs. Related notes: Microservices distributed locking, PostgreSQL row locks and concurrent updates and Redis rate limiting with Lua for more atomic scripts. Primary sources: Distributed Locks with Redis, the SET command reference, Martin Kleppmann’s How to do distributed locking and antirez’s reply, Is Redlock safe?.