In an offline-first React Native app using TanStack Query optimistic updates, why can a scan vanish from the screen and never sync, and when should rollback happen?
answer
- onMutate runs before any pause
- retry defaults to 0 for mutations
- network error is not a rejection
- roll back only on a final refusal
- undo this write, not a stale snapshot
basics
~20 sIf TanStack Query thinks it is online, an offline scan fails with a network error; with the default retry of 0, onError rolls it back and ends it. Pause offline, retry transient errors, roll back only on a definitive rejection.
solid answer
~40 s`onMutate` applies the optimistic scan before the request or any pause. If the `onlineManager` was not wired to NetInfo, or the phone is on a network with no route out, TanStack Query believes it is online, the request fails with a network error, and because mutations default to `retry: 0`, `onError` runs at once: the rollback removes the scan and the mutation ends in error, so nothing retries it. The fix is to wire the manager so offline writes pause, give mutations a `retry` function that retries network errors, timeouts, 5xx, 408 and 429, and treat `onError` as terminal: roll back only there, surface the failure instead of hiding it, and undo just this write's change rather than restoring an old snapshot that would erase later scans.
code
typescript · 24 linesimport { QueryClient } from '@tanstack/react-query';
import { isTransient, postScan, type PickLine, type Scan } from './api';
export const queryClient = new QueryClient();
const adjustPicked = (scan: Scan, delta: number) => (lines?: PickLine[]) =>
lines?.map(line =>
line.sku === scan.sku ? { ...line, picked: line.picked + delta } : line,
);
queryClient.setMutationDefaults(['recordScan'], {
mutationFn: (scan: Scan) => postScan(scan),
retry: (failureCount, error) => isTransient(error) && failureCount < 5,
onMutate: async (scan, context) => {
await context.client.cancelQueries({ queryKey: ['picklist', scan.orderId] });
context.client.setQueryData<PickLine[]>(['picklist', scan.orderId], adjustPicked(scan, +1));
},
onError: (_error, scan, _result, context) => {
// Terminal: undo only this scan, then surface it as failed in the UI.
context.client.setQueryData<PickLine[]>(['picklist', scan.orderId], adjustPicked(scan, -1));
},
onSettled: (_data, _error, scan, _result, context) =>
context.client.invalidateQueries({ queryKey: ['picklist', scan.orderId] }),
});go deeper
Know that an optimistic update shows a change before the server confirms it, and that a lost connection is not the same as the server saying no.
Explain why onMutate runs before the pause, why retry defaults to 0 for mutations, and how an unwired onlineManager turns an offline scan into an immediate onError.
Demonstrate the policy: pause offline, retry transient errors with a function, roll back only on terminal errors with a targeted undo, and surface pending and failed writes to the user.
Set the team's rule for which server responses are final, since the client's rollback policy and the API's error contract must agree across every offline feature.
## The symptom The warehouse app uses TanStack Query for its pick lists and applies each scan optimistically: the line's picked count goes up the moment the barcode is read. Pickers report that in some aisles a scan appears, then disappears a few seconds later, and never reaches the server. The mutation's `onError` restores the previous list, which is exactly what the optimistic-update pattern says to do. The problem is **when** `onError` runs, and what the app does there. ## How the scan disappears - `onMutate` runs as soon as `mutate()` is called, **before** the request is sent and before any pause, so the optimistic count shows immediately. - In these aisles the phone is attached to a Wi-Fi network that has no route out, or the `onlineManager` was never wired to NetInfo. Either way TanStack Query believes it is online, so the mutation does not pause; the request is attempted and fails with a network error. - Mutations default to `retry: 0`, so that first failure is final. `onError` runs, the rollback removes the scan from the screen, and the mutation ends in the `error` state. Nothing will retry it. The rollback did not fix anything: it turned a transient network problem into a lost write, and hid it. ## The rule: roll back only on a final, meaningful answer A queued write should be rolled back only when the system has **definitively** decided it will not apply: | Failure | Meaning | What the app should do | |---|---|---| | Offline (manager reports offline) | Not attempted yet | Keep the optimistic state; the mutation is paused | | Network error, timeout, 5xx, 408, 429 | Unknown or temporary | Keep it and retry, with backoff | | Validation error or other 4xx | The server refused this write | Roll back this write and tell the user | | Conflict (for example 409) | The server's state moved on | Roll back or merge, per the conflict policy | In TanStack Query terms: 1. **Wire the `onlineManager` to NetInfo** so writes made offline pause instead of failing. 2. **Give mutations a `retry` function** that returns `true` for transient errors, for example `(failureCount, error) => isTransient(error) && failureCount < 5`. Between attempts, a retrying mutation pauses again while the manager reports offline. 3. **Treat `onError` as terminal.** It runs once retries are exhausted or the error was not retryable. Roll back there, but surface the failure ("scan not sent, tap to retry") instead of silently removing it, so a transient error that outlasted the retries is not lost unseen. ## Roll back the write, not the world The common optimistic pattern snapshots the whole list in `onMutate` and restores that snapshot in `onError`. Offline, that is dangerous: while one scan waits, the picker makes ten more, each with its own optimistic change. Restoring the first scan's snapshot erases all ten. Prefer a **targeted undo** that reverses only this write's change (decrement this line, remove this item by its client id), and let `onSettled` invalidate the query so the server's truth replaces the optimistic state once the device is online. ## Why the snapshot pattern exists at all The snapshot-and-restore pattern comes from online apps, where a mutation usually settles within a second and nothing else changes the list in between, so the snapshot is still accurate when `onError` runs. Offline, a mutation can stay paused for an hour while the user keeps working, and the snapshot grows more wrong with every later write. The same code that was correct online becomes destructive offline, which is why interviewers ask about it for offline-first apps specifically. ## Showing pending state honestly - Mark optimistic entries as pending (`isPaused` from `useMutation`, or a flag in the cached data) so the picker can see what has not reached the server. - Keep a visible count of paused and failed writes. - Never let a rollback be the only signal that a write failed. ## What stays outside this question How to merge two devices' conflicting edits, and how the server should model versions, belong to offline-sync architecture. The client-side question is narrower: which errors mean "try again" and which mean "this will never apply", and making sure only the second kind undoes what the user saw.
- Why is restoring the snapshot taken in onMutate risky for a React Native app that queues many offline writes?Each snapshot is the list as it was before that one write. While a write waits offline, later writes change the same list optimistically. Restoring the first write's snapshot on failure erases every later change too, even though those writes may still succeed. Reversing only the failed write's own change, then invalidating on settle, keeps the other pending writes visible.
- In TanStack Query, what does a retrying mutation do between attempts while the device is offline?Before each new attempt the retryer checks whether it can continue; with the default `networkMode: 'online'` that requires the `onlineManager` to report online. If it does not, the mutation pauses and continues when connectivity returns, instead of burning its retries against a dead network. This only works if the manager is wired to NetInfo in React Native.
saying these in an interview costs you the question
- Every mutation error should roll back the optimistic update immediately.
- TanStack Query retries failed mutations three times by default.
- onError runs every time a single attempt fails, even when retries remain.
- Restoring the onMutate snapshot is always the safest rollback.
- A network error proves the server rejected the write.