skip to content

In a ride-hailing app, why are live driver locations usually held in memory with a short expiry instead of written durably on every update?

level: juniorimportance: must knowfreq 58%

answer

  1. value is obsolete within seconds
  2. only the latest position matters
  3. durability costs disk and replication
  4. expiry doubles as a liveness check
  5. index refills on the next pings

basics

~20 s

Each location ping is replaced by a newer one within seconds, so only the latest position matters. Keeping it in memory with a time-to-live makes writes cheap and removes silent drivers automatically once their updates stop.

solid answer

~40 s

A driver's location is **high-volume, short-lived data**: every ping supersedes the previous one a few seconds later. Writing each ping durably would spend disk I/O, replication and storage on values nobody will read again. So the live position is kept in an **in-memory store**, keyed by driver id, holding only the latest position. Each entry carries a **time-to-live (TTL)** of a few update intervals, for example 15-20 seconds when drivers report every 4 seconds. If the app crashes or loses signal, the entry expires and the driver stops being offered rides, with no cleanup job. Losing the store costs little: every active driver re-reports within one interval, so the index refills itself. Data that must survive, such as trips, assignments and billing, lives in a durable store.

go deeper

for a junior

Remember the core idea: a location ping is outdated a few seconds later, so the system keeps only the latest one in memory and lets it expire.

for a middle

Explain how the time-to-live works as a liveness signal and how to size it relative to the update interval so drivers neither flicker nor linger as ghosts.

for a senior

Show that you separate rebuildable state from state that must survive a restart, and describe how trip trails reach durable storage without slowing ingest.

for a principal

Frame the decision as cost of durability against the value of the data over time, and be explicit about which failures become brief degradations rather than data loss.

## What a location update is In a **ride-hailing app**, every online driver's phone sends its position to the backend on a fixed or adaptive interval, typically every few seconds. A single update is small: a driver id, latitude, longitude, a timestamp, and often heading, speed and status. The backend uses the **latest** update for one job: answering "which available drivers are near this pickup point right now?" The key property is that **each update makes the previous one obsolete**. A position from 30 seconds ago is not merely old; for matching it is wrong. ## Why not write every ping durably A durable database write is built to survive crashes: it goes to a write-ahead log, is flushed to disk and is often replicated before it is acknowledged. That machinery is worth paying for when the data must never be lost. For live positions it buys almost nothing: - **Volume** - with hundreds of thousands of drivers each reporting every few seconds, the system takes on the order of 10^5 writes per second. Durable storage at that rate needs many servers and constant compaction. - **Short value** - the stored value is overwritten a few seconds later, so the effort spent making it durable is wasted almost immediately. - **Latency** - matching reads positions constantly; memory serves those reads in microseconds. - **Storage growth** - keeping every ping forever turns a small working set into terabytes per day. - **Easy recovery** - if the in-memory store loses its data, active drivers re-report within one interval and the index rebuilds itself. ## What the in-memory record looks like The store keeps **one entry per driver**, overwritten on every update: ```json { "driverId": "d-18342", "lat": 40.7412, "lng": -73.9897, "heading": 270, "status": "AVAILABLE", "reportedAt": "2026-09-17T08:15:04Z", "ttlSeconds": 20 } ``` The working set stays small. Assuming roughly 200 bytes per entry, one million online drivers need about 200 MB of memory, plus overhead for the spatial index. ## Why the expiry matters The **time-to-live** is not just a way to save memory; it is the system's **liveness signal**. 1. A driver's app keeps sending updates, and each update resets the expiry. 2. The app crashes, the phone loses signal, or the driver drives into a tunnel. 3. No update arrives, so after the TTL the entry disappears. 4. The matcher no longer sees the driver and does not offer them rides they cannot answer. Choosing the TTL is a trade-off: - **Too short** (shorter than the update interval) - entries vanish between normal updates, and drivers flicker in and out of the index. - **Too long** (minutes or hours) - drivers who went offline keep being offered rides, which causes timeouts and wasted offers. - **About three to five update intervals** - tolerates a missed ping or two without showing ghost drivers for long. ## What still gets persisted Not everything is ephemeral. The split is by how long the data stays valuable: | Data | Where it lives | Why | |---|---|---| | Latest driver position | In-memory store with TTL | Needed for seconds, read constantly | | Driver availability and assignment state | Durable store with conditional writes | Must not be lost or double-assigned | | Trips, fares, payments | Durable database | Legal and financial record | | Route trail of a trip in progress | Appended asynchronously, often sampled or batched | Needed for receipts, disputes and analytics, not for matching | A common pattern is to write every ping to memory for matching and, **only for drivers on a trip**, also publish it to an append-only log that a background consumer stores in bulk. The durable path is then decoupled from the latency-sensitive matching path. ## Recovery and its limits Because the live index is reconstructible, a server failure is a short degradation, not data loss: for one update interval, drivers from that server are missing, then they reappear. This works only if the index holds **nothing that cannot be rebuilt from the next round of pings**. Anything that cannot be rebuilt, such as whether a driver has accepted a ride, must not live only in the ephemeral store. ## Common mistakes to avoid - Treating the location store as the **system of record** for anything but positions. - Resetting the expiry from the server side, which keeps silent drivers visible. - Assuming the in-memory store must be replicated as carefully as a payments database; a lost copy costs one update interval.

  • What should never be stored only in the ephemeral location store, and why?
    Anything that cannot be rebuilt from the next round of pings: whether a driver has accepted an offer, the trip record, payment state. If the in-memory store restarts, positions come back within one update interval, but an assignment would be lost, and the driver could be offered a second ride. That state belongs in a durable store that supports conditional writes.
  • How would you keep a route trail for trips without slowing down location ingest?
    Keep the hot path in memory, and for drivers on a trip also append each ping to an append-only log or queue. A background consumer writes those events to durable storage in batches, often downsampled. The ingest request never waits for durable storage, and a slow durable store only increases consumer lag.

It is like a whiteboard showing where each taxi is right now: you wipe and rewrite each entry constantly, and nobody files the old versions away.

saying these in an interview costs you the question

  • Every location ping should be written straight into the main relational database.
  • Keeping every historical ping is needed to find nearby drivers.
  • Offline drivers need a scheduled cleanup job to remove them from the index.
  • A time-to-live shorter than the update interval keeps the index fresher.
  • Losing the in-memory location store means permanent loss of driver data.