skip to content

In a chat app, why is a user's online status usually kept by periodic heartbeats with a TTL instead of connect and disconnect events?

level: juniorimportance: should knowfreq 56%

answer

  1. failures are silent
  2. half-open connections, suspended apps
  3. expiry as the default
  4. about three missed beats
  5. any device alive

basics

~20 s

Many disconnects never produce an event: a phone loses signal, an app is suspended, a gateway server crashes. With a heartbeat, the client refreshes a presence entry that expires unless renewed, so silence alone marks the user offline within a bounded time.

solid answer

~50 s

Presence has to be right when things **fail**, and failures are silent. A clean disconnect sends a close event. A dropped mobile connection, a suspended app, or a crashed gateway server sends nothing, so an event-only design leaves users "online" indefinitely. With heartbeats, the client (or its gateway) refreshes a presence entry every *N* seconds, and the entry carries a **TTL** of a few intervals, for example a 30 s heartbeat with a 90 s TTL. If the renewals stop, the entry expires and the user reads as offline within one TTL, with no cleanup code. A clean disconnect can still delete the entry at once, as a fast path. On expiry the server records **last seen**. A user with several devices is online while any device's entry is alive. The trade-off is staleness bounded by the TTL, plus a steady stream of heartbeat writes.

go deeper

for a junior

Remember the core idea: disconnects are often silent, so presence expires unless the client keeps renewing it with heartbeats.

for a middle

Explain how interval and TTL interact, why a TTL of a few intervals tolerates a lost beat, and how per-device entries give multi-device presence.

for a senior

Estimate the heartbeat write load, justify a separate in-memory tier with native expiry, and handle flapping with a grace period.

for a principal

Treat the interval as a product and cost decision: freshness against write volume and client battery, set per platform rather than globally.

## What presence means **Presence** is the "online" dot or "last seen 5 minutes ago" line in a chat app. It sounds trivial: mark the user online when they connect and offline when they disconnect. The difficulty is that the interesting cases are the ones where **no disconnect ever arrives**. ## Why connect and disconnect events are not enough - **Silent network loss**: a phone drives into a tunnel. The server's connection stays half-open until a transport timeout, which can take minutes. - **App suspension**: a mobile OS freezes a backgrounded app. The app sends no goodbye. - **Server crashes**: the gateway server holding 50,000 connections dies. No code runs to mark those 50,000 users offline. - **Lost events**: even when a disconnect event is emitted, it can be dropped on the way to the presence store. In every one of these cases, an event-only design leaves a **ghost**: a user who shows as online indefinitely. Fixing ghosts afterwards needs sweeper jobs that must themselves guess who is really there. ## How TTL heartbeats work 1. While connected, the client sends a small **heartbeat** every *N* seconds, for example every 30 s. The gateway can also send heartbeats on the client's behalf as long as the connection is healthy. 2. Each heartbeat **sets or refreshes** a presence entry with a **time-to-live (TTL)** a few intervals long, for example 90 s. That allows about three missed beats before the entry expires. 3. If the heartbeats stop for any reason, nobody has to *do* anything. The entry **expires** and the user reads as offline. 4. On a clean disconnect, the server may delete the entry right away. That is a **fast path**, not the only path. The design is **fail-safe**: the default outcome of silence is "offline", which is the truthful answer. ## Choosing the numbers | Setting | Shorter | Longer | |---|---|---| | Heartbeat interval | fresher presence, more writes and battery use | fewer writes, slower detection | | TTL (in intervals) | flickers offline on a brief blip | ghosts linger longer | A worked load estimate, assuming 10 million concurrently online users and one heartbeat every 30 s: 10,000,000 / 30 is about **333,000 presence writes per second**. That is why the presence store is usually a separate, in-memory key-value tier with native expiry, not the message store. The number also explains why intervals are rarely very short. ## Last seen, devices and flapping - **Last seen**: store the time of the last heartbeat, or the expiry time, so the app can show "last seen 10:42" once the user is offline. - **Multiple devices**: keep one entry per device. The user is online if **any** device's entry is alive, and offline only when all have expired. - **Flapping**: a user on a weak link can bounce between states. A short grace period before *announcing* offline avoids a storm of status changes for everyone watching. - **Privacy**: many products let users hide presence or last seen. The server must enforce that filter before sending presence to anyone else. ## What this does not cover How the client displays the dot, and how it spaces out reconnect attempts, belong to client-side design. How presence changes are fanned out to a large audience is its own cost problem. This mechanism only answers: **is this user actually here right now?**

  • How do you estimate the write load heartbeats create?
    Divide concurrently online users by the heartbeat interval. For example, 10 million users beating every 30 seconds is about 333,000 writes per second. That load is why presence usually lives in a dedicated in-memory store with native expiry, and why teams avoid very short intervals.
  • How should presence work when a user is signed in on a phone and a laptop?
    Keep a separate entry with its own TTL for each device. The user is online while any device's entry is alive, and becomes offline, with last seen recorded, only when the final entry expires. One device going idle does not hide the other.

It works like a lone hiker who must phone in every hour. Rescuers act when a check-in is missed. Nobody waits for the hiker to call and report that they fell.

saying these in an interview costs you the question

  • Set online on connect, offline on disconnect; nothing else is needed.
  • A clean disconnect event is always sent, so heartbeats are redundant.
  • Use a TTL equal to the heartbeat interval so presence is instant.
  • Store presence as a column in the main messages database.
  • A user with two devices goes offline when either device disconnects.