Background Sync is implemented only in Chromium browsers. For a field-service web app whose technicians submit reports from areas with poor connectivity, on Android and iOS alike, how would you design the offline write path?
answer
- the queue is the product
- triggers are the optimisation
- idempotency before reliability
- one flush, many callers
- show pending, never fake saved
basics
~20 sBuild a durable outbox in IndexedDB with server-side idempotency keys, and drive it from a ladder of triggers every browser has: immediate send, the online event, visibility changes, and app launch. Treat Background Sync as an extra trigger on Chromium, never the foundation.
solid answer
~50 sThe design principle is that the queue is the product and the trigger is an optimisation. Writes go into a durable outbox — IndexedDB, not localStorage — with a client-generated idempotency key, before the UI confirms anything. A flush routine drains it and is invoked from every trigger available: right after enqueue, on `online`, on `visibilitychange` when the page becomes visible, at app launch, and on Chromium additionally from a `sync` event so flushing can happen with no tab open. The server treats the idempotency key as the deduplication boundary, so replays are free. The UI must tell the truth: items show as pending until acknowledged, with a visible count and a manual retry. Then I would instrument outbox age and flush success, because on iOS the honest guarantee is weaker and the team needs to see how much weaker.
go deeper
Understand that offline support means storing the write locally first and sending it later, and that the user must be able to see something is still pending.
Explain the outbox plus flush-trigger pattern and why the same flush function should be driven by the online event, visibility changes and app launch rather than duplicated per trigger.
Bring idempotency keys and single-flight guarding into the design unprompted, and describe what the UI is allowed to claim about a queued item on a platform where flushing needs the app to be open.
Own the tradeoff explicitly: state the guarantee you can make per platform, instrument the gap so it is a number, define retention and dead-lettering for stale writes, and set the threshold at which the answer changes to requiring install or going native.
## Framing the decision The temptation is to compare APIs. The better framing is that Background Sync only changes *when* a flush is attempted; it never changes what must be true for the flush to be safe. So the architecture is the same everywhere, and platform support only determines how promptly the queue drains. That reframing is what a principal-level answer is expected to supply, because it turns an unfixable capability gap into a difference of degree. ## Layer 1 — the durable outbox A write is accepted into local storage before the network is touched. IndexedDB is the right store: it is asynchronous, holds structured records and blobs (field reports mean photos), and is not capped at a few megabytes of strings. Each entry carries the payload, an idempotency key generated at enqueue, a created timestamp, an attempt count, and a status. Two consequences follow immediately. First, the UI can only ever claim *queued*, never *saved*, until the server acknowledges. Second, the outbox is user data, so it inherits real obligations — it can contain personal information sitting on a shared device, and it is subject to storage eviction. That argues for requesting persistent storage where available and for a retention policy that discards or escalates entries beyond a certain age rather than retrying forever. ## Layer 2 — idempotency at the server Because every trigger may fire while another attempt is in flight, and because a response can be lost after the server committed, the server must be able to recognise a repeat. A client-generated key on the record, honoured server-side for a defined window, converts an at-least-once delivery pipeline into an effectively-once one. Without it, the more reliable your retry triggers are, the more duplicates you create — the failure mode gets worse as the client gets better, which is why this layer comes before any discussion of triggers. ## Layer 3 — the trigger ladder One flush function, many callers: ```js async function flushOutbox() { if (flushOutbox.running) return flushOutbox.running; // single-flight flushOutbox.running = drain().finally(() => (flushOutbox.running = null)); return flushOutbox.running; } window.addEventListener('online', flushOutbox); document.addEventListener('visibilitychange', () => { if (document.visibilityState === 'visible') flushOutbox(); }); window.addEventListener('load', flushOutbox); ``` And on Chromium, additionally, a `sync` registration so the same drain runs in the service worker with no page open. Every trigger calls the same code path, guarded by single-flight so overlapping triggers cannot double-send. What this ladder buys per platform is worth stating plainly in an interview: on Chromium a queued report can leave the device while the app is closed; on iOS Safari it leaves the next time the user opens the app with signal. For technicians who open the app several times a day, that difference is hours, not days — and knowing that is what decides whether the gap is acceptable rather than a reason to go native. ## Layer 4 — honest UI The most common real-world failure of offline apps is not lost data; it is a green checkmark shown for data that never left the device. The queue must be visible: a pending count, per-item state, a manual *retry now*, and an explicit warning before the user signs out or clears data with items outstanding. On iOS the app should nudge — "3 reports waiting to upload, keep the app open a moment" — because the user is the trigger there. ## Layer 5 — measurement Instrument time-from-enqueue-to-acknowledged as a distribution, per platform; the count of entries older than a threshold; flush attempts per success; and eviction events. Without this, the iOS/Chromium difference is an argument in a meeting rather than a number. With it, the team can decide whether pushing users to install the app — which on iOS also unlocks push from iOS 16.4 — is worth the friction. ## What I would not do I would not build two code paths, one Background-Sync-shaped and one not; the divergence rots. I would not put pending writes in `localStorage`, which is synchronous, small and string-only. I would not retry indefinitely without a ceiling, because a permanently poisoned record then blocks the queue forever — failed items move to a dead-letter state visible to the user and to support. And I would not treat the offline path as a feature to add later: the outbox has to be the only write path from day one, or half the codebase will keep calling `fetch` directly and the guarantee will be a fiction.
- Why does adding more reliable retry triggers make server-side idempotency more urgent rather than less?Every additional trigger increases the number of attempts, and an attempt whose response is lost after the server committed is indistinguishable from a failure. So the better the client gets at retrying, the more duplicate side effects it produces. A client-generated idempotency key honoured server-side is what makes at-least-once delivery safe.
- Why IndexedDB rather than localStorage for the outbox?localStorage is synchronous, string-only, and capped at a few megabytes per origin, so it blocks the main thread and cannot hold photo attachments. IndexedDB is asynchronous, stores structured values and blobs, and has a far larger quota — the practical requirements of an outbox that must survive a tab close.
- What would make you conclude the iOS gap is unacceptable and change the plan?Measured data: if the distribution of enqueue-to-acknowledged time on iOS shows reports routinely sitting for days, or eviction is dropping entries, the gap is real harm. The responses in order of cost are nudging install, then a hard requirement to sync before sign-off, then a native wrapper — chosen on the numbers, not on principle.
- How do you stop one permanently failing record from blocking the whole queue?Cap attempts and move exhausted entries to a dead-letter state instead of leaving them at the head of the queue. Keep draining the rest, surface the failed item to the user with its reason, and expose it to support. Unbounded retries on a poisoned record turn a single bad write into a total outage of the write path.
saying these in an interview costs you the question
- Building the whole design around Background Sync and bolting on a fallback
- Confirming a write in the UI before the server has acknowledged it
- Queueing pending writes in localStorage
- Retrying failed items forever with no dead-letter state
- Assuming an installed PWA gets background execution on iOS