In a ride-hailing app, how does the driver location update interval trade position accuracy against write load on the backend?
answer
- two linear formulas
- load scales with 1 / interval
- staleness scales with speed
- phone positioning error sets a floor
- vary rate by driver state
basics
~20 sWrite rate is drivers divided by interval, while staleness is speed times interval. One million drivers every 4 seconds means 250,000 writes per second, and a driver at 50 km/h may be about 56 m from the stored point.
solid answer
~40 sTwo formulas drive the choice. **Write rate = active drivers / interval**, and **worst-case staleness = speed x interval**. With one million drivers reporting every 4 seconds, the backend takes 250,000 writes per second; at 50 km/h (about 13.9 m/s) the stored position can be about 56 m behind. Halving the interval halves the staleness but doubles the load, battery use and mobile data. In practice, the interval is **adaptive**: frequent while a driver is heading to a pickup or on a trip, slower when idle or parked, and often triggered by distance moved rather than by time alone. Below a certain interval, extra pings add little, because the phone's own positioning error is already several metres or worse among tall buildings.
code
pseudocode · 12 linesfunction shouldSendUpdate(state, now, lastSentAt, movedMeters):
if state == OFFLINE:
return false
if state == EN_ROUTE or state == ON_TRIP:
maxGap = 4 seconds
minMove = 20 meters
else: # AVAILABLE and idle
maxGap = 15 seconds
minMove = 100 meters
if now - lastSentAt >= maxGap:
return true # heartbeat keeps the server-side expiry alive
return movedMeters >= minMovego deeper
Know that more frequent updates mean fresher positions but more writes, and be able to divide drivers by interval to get writes per second.
Work the arithmetic out loud: writes per second, staleness at a stated speed, and bytes per day, stating each assumption along the way.
Show adaptive policies by state and movement, ordering by client timestamp, and why shedding stale updates beats queueing them during overload.
Balance server cost, driver battery and data use, and pickup accuracy, and name where the phone's own positioning error makes faster updates pointless.
## The two formulas A **ride-hailing app** asks each online driver's phone to report its position every **T** seconds. Two quantities follow directly from T: - **Write rate** = number of active drivers / T. This is the steady load the ingest tier must absorb. - **Worst-case staleness** = driver speed x T. Just before the next update, the stored position can be this far behind the real one. Everything else - servers, bandwidth, battery - scales with the write rate, and matching quality depends on staleness. ## Worked numbers Assume **one million** online drivers, a typical urban speed of **50 km/h** (50,000 m / 3,600 s, about 13.9 m/s) and roughly **100 bytes** per update on the wire. These are illustrative figures, not a real service's numbers. | Interval | Writes per second | Max staleness at 50 km/h | Raw bytes per day if kept | |---|---|---|---| | 1 s | 1,000,000 | about 14 m | about 8.6 TB | | 4 s | 250,000 | about 56 m | about 2.16 TB | | 10 s | 100,000 | about 139 m | about 0.86 TB | Checks: 13.9 x 4 = 55.6 m; 250,000 x 100 bytes = 25 MB/s; 25 MB/s x 86,400 s = 2,160,000 MB, about 2.16 TB. The last column is why the live index keeps only the latest point and does not store the raw stream for matching. ## What accuracy is actually needed Staleness matters less than it first appears: - **Matching** only has to rank nearby drivers correctly. If every candidate is 50 m stale, the ranking by pickup time rarely changes. - **The rider's map** can smooth movement between updates by interpolating along the road, so a slower interval still looks fluid. - **Positioning error** on a phone is typically a few metres under open sky and worse among tall buildings. Once the interval's staleness is close to that error, faster updates mostly add noise. - **Arrival detection** ("your driver is here") is where freshness matters most, which is why intervals shrink near the pickup. ## Adaptive intervals A fixed fast interval for everyone wastes most of its writes on drivers who are parked. Common refinements: 1. **By state** - frequent updates while driving to a pickup or on a trip; slower while idle and available; none while offline. 2. **By movement** - send when the driver has moved more than a distance threshold, or when a maximum time has passed, whichever comes first. A parked driver then sends only occasional heartbeats, and those heartbeats keep the entry's time-to-live alive. 3. **By demand** - in areas with few riders waiting, idle positions can be coarser. 4. **Client-side batching** - on a trip, the phone can collect several positions and upload them together. The trail stays detailed while the number of requests stays low. If most online drivers are idle at any moment, state-based intervals cut the write rate sharply without hurting pickups. ## The costs beyond the server The interval is also a **client** decision: - **Battery** - frequent positioning and radio use drain the driver's phone during a long shift. - **Mobile data** - each update costs bytes; a persistent connection such as a WebSocket avoids repeating connection setup and headers on every update. - **Network variance** - updates arrive late or in bursts on poor connections, so the server should order them by the phone's timestamp and drop an update older than the stored one. ## Designing the ingest tier With the write rate known, the ingest path is sized to it: - keep the handler tiny: validate, compute the region or cell, overwrite the in-memory entry, reset the expiry; - partition ingest by region so load spreads and each partition holds a bounded number of drivers; - plan for peaks (evening rush, events), not the daily average, and add headroom; - when overloaded, prefer **dropping or coalescing** stale updates over queueing them, because a queued old position is worth less than the next fresh one. That last rule is specific to this workload: unlike orders or payments, a delayed location update loses value every second it waits. ## Choosing a number A reasonable starting point is to pick the **staleness you can tolerate** at pickup (tens of metres), divide it by a typical urban speed to get the interval, and then check the resulting write rate against your ingest budget at peak. If the budget is exceeded, adapt by state before you shorten anything else.
- Updates arrive out of order on a flaky mobile connection. How should the ingest handler react?Store each update with the phone's own timestamp and overwrite the entry only if the incoming timestamp is newer than the stored one. An older update that arrives late is dropped, so a delayed packet cannot move the driver backwards on the map or into the wrong region.
- Under an ingest overload, should you queue location updates until capacity recovers?Usually not. A location update loses value every second, and the next one from the same driver replaces it. Coalescing per driver or shedding the oldest updates keeps the index as fresh as possible, while a long queue would spend capacity writing positions that are already out of date.
saying these in an interview costs you the question
- Sending updates every second is always better because positions are fresher.
- Doubling the interval doubles the write rate.
- The raw ping stream must be stored in full for matching to work.
- Every online driver should report at the same fixed rate.
- Queued old location updates must all be applied in order during overload.