A case repository refreshes its stored tracker links only when the tracker calls it back on a change. Why does it still drift out of agreement over time, and what does a periodic reconciliation sweep add?
answer
- you cannot notice a silence
- some changes announce nothing
- changed-since never lists what is gone
- walk your own links to find absence
- a busy sweep means a broken fast path
basics
~20 sCallbacks are best-effort and only fire for changes a tracker chooses to announce, so a lost delivery leaves the repository confidently wrong with no visible signature. A periodic sweep re-reads stored links and is the only path that finds items that quietly went missing.
solid answer
~50 sA callback-driven mirror has no way to notice what it did not receive. A delivery dropped during a restart or a credential expiry is simply never seen again, and nothing in the stored state says so — the stale value looks exactly like a fresh one. Some changes also announce nothing at all: an item leaving the integration account's visibility, a bulk administrative edit, a project archived. So the callback stream is the fast path and a **reconciliation sweep** is the correctness path. A cheap sweep asks the tracker for everything modified since a stored watermark and applies it, which catches missed updates but never catches a disappearance. Only re-resolving every stored link, on a longer cadence, finds links whose target is gone. Both routes should end in the same apply function so they cannot disagree.
go deeper
Know that a callback tells you about a change quickly but is not guaranteed to arrive, so a periodic re-read is what actually keeps two systems agreeing.
Explain the difference between asking for everything changed since a stored timestamp and re-resolving every stored link, and say which of the two can find a deleted item.
Show the operational discipline: a watermark advanced only after durable writes, slightly overlapping windows, batched and paced reads, resumable runs, and one shared apply path for both routes.
Own the tradeoff between sweep cadence, tolerable drift and the load you place on a tracker other teams depend on, and be able to justify the interval you chose to the people who operate it.
A tracker-side callback is a performance optimization. It tells you about a change in seconds instead of at the next sweep. It is not, and cannot be, the thing that makes two stores agree, because a consumer has no way to notice a message it never received. ## What a callback stream cannot tell you Three distinct gaps, and they fail differently: - **The delivery you never got.** A restart, an expired credential, a lapsed certificate, a receiver that returned errors until the sender gave up. The change happened, the announcement did not arrive, and the stored value is now wrong while looking exactly as trustworthy as one updated a second ago. Staleness has no visible signature. - **The change that announces nothing you subscribed to.** Visibility changes are the classic case: nobody edited the item, but the integration account lost access to its project, so every link into it now resolves to a not-found. Bulk administrative operations and project archival often behave the same way. - **Order.** Announcements are not guaranteed to arrive in the order the changes happened. Applying an older state after a newer one leaves the repository stably wrong until the next change comes along. None of these produce an error on the receiving side, so a callback-only design fails silently and indefinitely. The first person to notice is usually a human reading a link whose state has not matched the ticket for a fortnight. ## Two sweeps, doing two different jobs A **reconciliation sweep** re-derives the repository's copy from the tracker instead of waiting to be told. It comes in two shapes, and using one where you needed the other is the common mistake: 1. **Watermark sweep (incremental).** Ask the tracker for everything changed since the timestamp of the last successful sweep, and apply it. Cost is proportional to churn, so it can run often. It catches every missed update — including ones you never knew you missed — and catches nothing about items that were deleted or that left your visibility, because those do not appear in a list of changes. 2. **Full sweep (resolve every stored link).** Walk your own links and resolve each identifier against the tracker. Cost is proportional to how many links you store rather than to churn, so it runs on a long cadence. This is the only pass that can discover **absence**: the link whose target no longer answers. The watermark itself is the subtle part. Store the tracker's notion of modification time, not your clock, and advance it only after the sweep's writes have committed — a watermark advanced before the work is durable turns one crash into a permanent hole. Overlapping the window slightly on each run is cheap insurance against boundary rounding, and is harmless as long as applying the same change twice is a no-op. ## Both paths must end in the same code The callback handler and the sweep must converge on identical state from identical input. The reliable way to guarantee that is structural: neither applies anything itself. Both produce "the tracker says this item now looks like this", and one shared apply function decides what changes locally. That makes every apply **idempotent** — running it twice changes nothing the second time — which is what lets the sweep overlap the callback stream freely, resume after a partial failure, and be re-run on demand after an incident without anyone having to reason about what might get applied twice. ## A sweep has a cost, and needs a budget A full sweep is the most expensive thing the integration does, and it lands on a system shared with everyone else in the company. The mechanics that make it acceptable are ordinary and non-negotiable: - Read in **batches** rather than one item per request, so the cost scales with batch count instead of link count. - **Chunk and pace** the work, back off exponentially when the tracker signals throttling, and treat throttling as an instruction to slow down rather than an error to retry immediately. - Make it **resumable**, so a sweep interrupted midway continues instead of starting over. - Run it **off-peak**, and treat its cadence as a configurable operational choice, because the right interval depends on how much drift the organization can tolerate. ## The sweep is also the only drift meter you have Instrument how many links each sweep actually changed. In a healthy integration that number sits near zero: the callbacks got there first, and the sweep merely confirms what you already believed. A sweep that keeps finding real changes is telling you the fast path is broken — an endpoint unreachable, a subscription silently dropped, an account whose access was narrowed. That signal exists nowhere else, which is why the sweep earns its cost even when it is usually a no-op.
- How would you choose the interval between full sweeps?Work backwards from the damage a stale link does: if it only misleads a weekly review, a sweep finishing before that review is enough. Then check feasibility — the sweep must complete comfortably inside its own interval on your largest link set, paced so the tracker keeps headroom for its human users. If those two constraints never overlap, store fewer links rather than sweeping less honestly.
- The sweep and a callback apply a change to the same link at the same moment. What stops them fighting?Make the apply idempotent and ordered. Both routes hand one function a version of the item plus its modification time, and the function ignores anything not newer than what it already stored. Two concurrent applies of the same state then produce one write and one no-op in either order, needing nothing heavier than a per-link compare-and-set.
The callback stream is a shop's till receipts and the sweep is its stocktake. You can post every receipt perfectly and still be wrong about what is on the shelves, because breakage and theft never write a receipt — the only way to find those is to go and count.
saying these in an interview costs you the question
- Treating a callback stream as a guaranteed record of changes
- Advancing the watermark before the sweep's writes commit
- Expecting a changed-since query to reveal deletions
- Running a full sweep as often as the incremental one
- Letting the callback path and the sweep apply changes differently