Ch. 19 · System Design

System Design: Presence Tracking

Track online status with heartbeats and TTLs, derive presence instead of storing it, and avoid fanning out to every client.

~2 min readadvancedupdated Oct 5, 2026

Presence answers “is this user online”. It is hard not because the state is complex but because it changes constantly and must stay cheap at scale. The standard approach is a heartbeat with a short TTL, so presence is derived from freshness rather than explicitly set to offline.

Before you start

You should understand TTLs, WebSockets and fan-out. This article is conceptual with small sketches.

Step-by-step walkthrough

Step 1: Write a heartbeat with a TTL

On each connection or activity, write a key presence:{userId} with a short TTL such as 30 seconds. The client refreshes it periodically. There is no explicit offline write; when the heartbeat stops, the key expires and the user is considered offline.

Step 2: Derive presence on read

A read checks whether the key exists. This makes a crashed client or a dropped network look offline without a cleanup job, and it tolerates clock and message delays as long as the TTL is longer than the heartbeat interval.

Step 3: Fan out selectively

Broadcasting every status change to every client is O(users squared) and does not scale. Send presence only for the contacts a client is currently viewing, subscribe per conversation, and batch updates. Consider a coarse status such as “active recently” instead of live precision for large audiences.

Worked scenario

The write is a heartbeat and the read is an existence check.

on activity:  SET presence:42 EX 30        # refresh TTL
is online?:   EXISTS presence:42           # derived, no offline write
after 30s idle without heartbeat: key expires -> offline
Text

Walk through the example

The client refreshes the key while active, so the TTL never expires. When the client disconnects and stops refreshing, the key lapses and the next check reports offline. No background job flips the status, which removes a source of inconsistency and load.

Common mistake

Storing a durable isOnline flag and updating it on disconnect, which breaks when a client crashes and never sends the disconnect. Another is broadcasting every change to all users, which explodes at scale.

Verify the behavior

Stop sending heartbeats and confirm the user reads as offline after the TTL. Simulate a crash with no disconnect message and confirm presence still lapses. Measure fan-out with a large audience and confirm subscriptions keep messages bounded.

Interview exercise

Why derive presence from a TTL instead of writing online and offline explicitly?

Answer and reasoning

A TTL makes the failure mode self-healing: a crashed or disconnected client simply stops refreshing and expires, without depending on a clean disconnect message that may never arrive. Explicit writes require reliable cleanup and create inconsistency when one is missed. The TTL trades exactness for robustness, which is the right call for presence.

Continue learning

Compare real-time patterns in WebSocket lifecycle and Feed fan-out. Read the Redis TTL documentation and try the System design interview questions.

More in System Design

read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
read ✓System Design · hard

System Design: Idempotent APIs

Make retried requests safe with idempotency keys, store the result per key, and return the original response on a repeat.

~2 min readread →
esc