skip to content

questions

11

In a ride-hailing app, why are live driver locations usually held in memory with a short expiry instead of written durably on every update?

level: juniorimportance: must knowfreq 58%

answer

  1. value is obsolete within seconds
  2. only the latest position matters
  3. durability costs disk and replication
  4. expiry doubles as a liveness check
  5. index refills on the next pings

basics

~20 s

Each location ping is replaced by a newer one within seconds, so only the latest position matters. Keeping it in memory with a time-to-live makes writes cheap and removes silent drivers automatically once their updates stop.

solid answer

~40 s

A driver's location is **high-volume, short-lived data**: every ping supersedes the previous one a few seconds later. Writing each ping durably would spend disk I/O, replication and storage on values nobody will read again. So the live position is kept in an **in-memory store**, keyed by driver id, holding only the latest position. Each entry carries a **time-to-live (TTL)** of a few update intervals, for example 15-20 seconds when drivers report every 4 seconds. If the app crashes or loses signal, the entry expires and the driver stops being offered rides, with no cleanup job. Losing the store costs little: every active driver re-reports within one interval, so the index refills itself. Data that must survive, such as trips, assignments and billing, lives in a durable store.

go deeper

for a junior

Remember the core idea: a location ping is outdated a few seconds later, so the system keeps only the latest one in memory and lets it expire.

for a middle

Explain how the time-to-live works as a liveness signal and how to size it relative to the update interval so drivers neither flicker nor linger as ghosts.

for a senior

Show that you separate rebuildable state from state that must survive a restart, and describe how trip trails reach durable storage without slowing ingest.

for a principal

Frame the decision as cost of durability against the value of the data over time, and be explicit about which failures become brief degradations rather than data loss.

## What a location update is In a **ride-hailing app**, every online driver's phone sends its position to the backend on a fixed or adaptive interval, typically every few seconds. A single update is small: a driver id, latitude, longitude, a timestamp, and often heading, speed and status. The backend uses the **latest** update for one job: answering "which available drivers are near this pickup point right now?" The key property is that **each update makes the previous one obsolete**. A position from 30 seconds ago is not merely old; for matching it is wrong. ## Why not write every ping durably A durable database write is built to survive crashes: it goes to a write-ahead log, is flushed to disk and is often replicated before it is acknowledged. That machinery is worth paying for when the data must never be lost. For live positions it buys almost nothing: - **Volume** - with hundreds of thousands of drivers each reporting every few seconds, the system takes on the order of 10^5 writes per second. Durable storage at that rate needs many servers and constant compaction. - **Short value** - the stored value is overwritten a few seconds later, so the effort spent making it durable is wasted almost immediately. - **Latency** - matching reads positions constantly; memory serves those reads in microseconds. - **Storage growth** - keeping every ping forever turns a small working set into terabytes per day. - **Easy recovery** - if the in-memory store loses its data, active drivers re-report within one interval and the index rebuilds itself. ## What the in-memory record looks like The store keeps **one entry per driver**, overwritten on every update: ```json { "driverId": "d-18342", "lat": 40.7412, "lng": -73.9897, "heading": 270, "status": "AVAILABLE", "reportedAt": "2026-09-17T08:15:04Z", "ttlSeconds": 20 } ``` The working set stays small. Assuming roughly 200 bytes per entry, one million online drivers need about 200 MB of memory, plus overhead for the spatial index. ## Why the expiry matters The **time-to-live** is not just a way to save memory; it is the system's **liveness signal**. 1. A driver's app keeps sending updates, and each update resets the expiry. 2. The app crashes, the phone loses signal, or the driver drives into a tunnel. 3. No update arrives, so after the TTL the entry disappears. 4. The matcher no longer sees the driver and does not offer them rides they cannot answer. Choosing the TTL is a trade-off: - **Too short** (shorter than the update interval) - entries vanish between normal updates, and drivers flicker in and out of the index. - **Too long** (minutes or hours) - drivers who went offline keep being offered rides, which causes timeouts and wasted offers. - **About three to five update intervals** - tolerates a missed ping or two without showing ghost drivers for long. ## What still gets persisted Not everything is ephemeral. The split is by how long the data stays valuable: | Data | Where it lives | Why | |---|---|---| | Latest driver position | In-memory store with TTL | Needed for seconds, read constantly | | Driver availability and assignment state | Durable store with conditional writes | Must not be lost or double-assigned | | Trips, fares, payments | Durable database | Legal and financial record | | Route trail of a trip in progress | Appended asynchronously, often sampled or batched | Needed for receipts, disputes and analytics, not for matching | A common pattern is to write every ping to memory for matching and, **only for drivers on a trip**, also publish it to an append-only log that a background consumer stores in bulk. The durable path is then decoupled from the latency-sensitive matching path. ## Recovery and its limits Because the live index is reconstructible, a server failure is a short degradation, not data loss: for one update interval, drivers from that server are missing, then they reappear. This works only if the index holds **nothing that cannot be rebuilt from the next round of pings**. Anything that cannot be rebuilt, such as whether a driver has accepted a ride, must not live only in the ephemeral store. ## Common mistakes to avoid - Treating the location store as the **system of record** for anything but positions. - Resetting the expiry from the server side, which keeps silent drivers visible. - Assuming the in-memory store must be replicated as carefully as a payments database; a lost copy costs one update interval.

  • What should never be stored only in the ephemeral location store, and why?
    Anything that cannot be rebuilt from the next round of pings: whether a driver has accepted an offer, the trip record, payment state. If the in-memory store restarts, positions come back within one update interval, but an assignment would be lost, and the driver could be offered a second ride. That state belongs in a durable store that supports conditional writes.
  • How would you keep a route trail for trips without slowing down location ingest?
    Keep the hot path in memory, and for drivers on a trip also append each ping to an append-only log or queue. A background consumer writes those events to durable storage in batches, often downsampled. The ingest request never waits for durable storage, and a slow durable store only increases consumer lag.

It is like a whiteboard showing where each taxi is right now: you wipe and rewrite each entry constantly, and nobody files the old versions away.

saying these in an interview costs you the question

  • Every location ping should be written straight into the main relational database.
  • Keeping every historical ping is needed to find nearby drivers.
  • Offline drivers need a scheduled cleanup job to remove them from the index.
  • A time-to-live shorter than the update interval keeps the index fresher.
  • Losing the in-memory location store means permanent loss of driver data.
open as a page

In a find-nearby-places service, what is a geohash, and how does it let an ordinary sorted index answer proximity queries?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A geohash turns latitude and longitude into one short string by interleaving their bits, so points in the same cell share a prefix. A prefix range scan on an ordinary sorted index then returns every point in that cell.

open as a page

In a ride-hailing app, how does the driver location update interval trade position accuracy against write load on the backend?

level: middleimportance: must knowfreq 62%

basics

~20 s

Write rate is drivers divided by interval, while staleness is speed times interval. One million drivers every 4 seconds means 250,000 writes per second, and a driver at 50 km/h may be about 56 m from the stored point.

open as a page

In a geohash-indexed nearby-places service, how do you choose the cell precision and scan pattern so a radius query misses no point?

level: middleimportance: must knowfreq 62%

basics

~10 s

Pick the finest precision whose cells are at least as large as the radius on their shorter side, scan the query's cell plus its eight neighbours, then keep only candidates within the exact distance.

open as a page

In a ride-hailing dispatch system with several concurrent matchers, how do you guarantee a driver is never assigned to two riders at once?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Keep each driver's dispatch state in an authoritative store and move it from available to offered with a conditional write that only one matcher can win. Offers expire, and accepting is another conditional write checked against the offer id.

open as a page

In a ride-hailing dispatch flow, how should candidate drivers found by a geo index be ranked using a separate ETA service?

level: middleimportance: should knowfreq 45%

basics

~20 s

Use the geo index only to shortlist a bounded set of nearby available drivers, then ask the ETA service for their road travel times in one batched call and rank by ETA, falling back to a distance estimate if it is slow.

open as a page

In a find-nearby-places service covering dense cities and empty oceans, how does a quadtree adapt its cells compared with a fixed-size grid?

level: middleimportance: should knowfreq 55%

basics

~20 s

A quadtree splits a region into four quadrants whenever it holds more than a set number of points, so dense areas get many small cells and sparse areas keep a few large ones, bounding the points per cell.

open as a page

In a cell-indexed nearby-places service, how do you answer 'the 10 nearest restaurants' correctly when the request gives no search radius?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Use an expanding ring search: scan the user's cell and then successive rings of neighbours, keeping the best k by exact distance, and stop only when the k-th best lies within the radius the scanned cells fully cover.

open as a page

In a ride-hailing dispatcher, when would you match riders to drivers in short batching windows instead of greedily assigning each rider the nearest driver?

level: principalimportance: should knowfreq 40%

basics

~20 s

Batch where demand and supply are dense enough that riders in the same few seconds compete for the same drivers; a joint assignment then lowers total pickup time. Where requests are sparse, greedy matching is simpler and adds no wait.

open as a page

In a ride-hailing app that shards its in-memory driver location index by region, what problems do shard boundaries cause and how are they handled?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Moving drivers leave stale entries in the old shard, searches near an edge miss drivers across it, and drivers on a border flip between shards. Fixes are handoff plus dedupe, fan-out to every overlapping shard, and hysteresis.

open as a page

Why might a proximity service index points with S2 or H3 cells instead of geohash rectangles when working on a spherical Earth?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Geohash cells come from a flat lat/lng grid, so they distort with latitude and jump along a Z-order curve. S2 cube-face cells with Hilbert ordering and hexagonal H3 cells keep sizes more uniform and neighbours more regular.

open as a page