A distributed lock ensures only one instance runs a critical section at a time, such as a scheduled job or a partition owner. It sounds simple but is easy to get wrong: a lock without an expiry deadlocks, and one without fencing lets a paused holder write after losing the lock.
Before you start
You should understand leases, timeouts and idempotency. This article is conceptual with a small illustration.
Step-by-step walkthrough
Step 1: Use a lease with an expiry
Acquire the lock as a key with a TTL and renew it while working, for example SET key owner NX PX 30000. The expiry ensures a crashed holder releases the lock automatically. A lock with no expiry is a future outage.
Step 2: Add a fencing token
Renewing or re-acquiring should increment a token, and the protected resource should reject writes with a token lower than the highest seen. This prevents a paused holder from acting after its lease expired, which is the failure a lease alone does not cover.
Step 3: Prefer structure over locks
Many lock use cases disappear with partitioning (each instance owns a subset so there is no contention) or idempotency (a duplicate run is harmless). Reach for a lock only when a genuine mutual-exclusion requirement remains, and keep the critical section short.
Worked scenario
An exclusive job holds a lease and fences its writes.
acquire: SET job:report owner=A NX PX 30000 -> ok (token A)
renew: periodic refresh while running
crash: key expires, owner=B acquires with a later token
stale A: writes carrying A's token are rejectedWalk through the example
The lease and its PX expiry let a crashed holder release automatically, and the token ordering rejects a stale writer. Without the expiry, a crash would leave the lock held forever; without the token, a paused holder could still write. Together they make the lock safe.
Common mistake
Holding a lock with no expiry, which deadlocks on a crash, or relying on a lock without fencing, which allows stale writes. Another is a long critical section, which increases contention and the chance of expiry mid-work.
Verify the behavior
Crash the holder and confirm the lease expires and another instance acquires it. Pause the holder past expiry and confirm its token is rejected. Assert only one instance runs the critical section at a time under load.
Interview exercise
When can you avoid a distributed lock entirely?
Answer and reasoning
When you can partition the work so each instance owns a disjoint subset and there is no shared resource to guard, or when the operation is idempotent so duplicates are harmless. Both remove the need for mutual exclusion. Locks are a last resort because they add coordination, latency and failure modes.
Continue learning
Compare job coordination in Concurrent job claiming and leader election in Leader election. Read the Redis distributed locks guidance and try the Microservices interview questions.