Presence answers “is this user online”. It is hard not because the state is complex but because it changes constantly and must stay cheap at scale. The standard approach is a heartbeat with a short TTL, so presence is derived from freshness rather than explicitly set to offline.
Before you start
You should understand TTLs, WebSockets and fan-out. This article is conceptual with small sketches.
Step-by-step walkthrough
Step 1: Write a heartbeat with a TTL
On each connection or activity, write a key presence:{userId} with a short TTL such as 30 seconds. The client refreshes it periodically. There is no explicit offline write; when the heartbeat stops, the key expires and the user is considered offline.
Step 2: Derive presence on read
A read checks whether the key exists. This makes a crashed client or a dropped network look offline without a cleanup job, and it tolerates clock and message delays as long as the TTL is longer than the heartbeat interval.
Step 3: Fan out selectively
Broadcasting every status change to every client is O(users squared) and does not scale. Send presence only for the contacts a client is currently viewing, subscribe per conversation, and batch updates. Consider a coarse status such as “active recently” instead of live precision for large audiences.
Worked scenario
The write is a heartbeat and the read is an existence check.
on activity: SET presence:42 EX 30 # refresh TTL
is online?: EXISTS presence:42 # derived, no offline write
after 30s idle without heartbeat: key expires -> offlineWalk through the example
The client refreshes the key while active, so the TTL never expires. When the client disconnects and stops refreshing, the key lapses and the next check reports offline. No background job flips the status, which removes a source of inconsistency and load.
Common mistake
Storing a durable isOnline flag and updating it on disconnect, which breaks when a client crashes and never sends the disconnect. Another is broadcasting every change to all users, which explodes at scale.
Verify the behavior
Stop sending heartbeats and confirm the user reads as offline after the TTL. Simulate a crash with no disconnect message and confirm presence still lapses. Measure fan-out with a large audience and confirm subscriptions keep messages bounded.
Interview exercise
Why derive presence from a TTL instead of writing online and offline explicitly?
Answer and reasoning
A TTL makes the failure mode self-healing: a crashed or disconnected client simply stops refreshing and expires, without depending on a clean disconnect message that may never arrive. Explicit writes require reliable cleanup and create inconsistency when one is missed. The TTL trades exactness for robustness, which is the right call for presence.
Continue learning
Compare real-time patterns in WebSocket lifecycle and Feed fan-out. Read the Redis TTL documentation and try the System design interview questions.