A sync creates an item in an issue tracker, then crashes before it stores the identifier the tracker returned, and on the next pass it creates a second item. How do you make that write safe to retry?
answer
- three outcomes, not two
- the response was what got lost
- record the attempt, not the result
- carry a token the item keeps
- search before you create again
basics
~20 sRecord the intent before calling the tracker and carry a locally generated token into the created item, so a retry searches for the token and adopts what the lost call already made. Creation is the one sync write that is not naturally repeatable.
solid answer
~50 sThe thing that was lost is the acknowledgement, not the write: the item exists, the caller never learned its identifier, and every later operation the sync wants to perform has nothing to address. Retrying the create is what turns one invisible failure into two visible tickets. Two mechanisms fix it. First, **write the intent down first** — persist a pending record carrying a token you generate before the call, so a crash leaves durable evidence that a create may be in flight. Second, **make the created item findable** — put that token somewhere durable and searchable on the item, so the retry searches for it before creating and adopts what it finds. Where the tracker accepts a caller-supplied idempotency token natively, prefer that: it does the same job server-side, where there is no race for the caller to lose.
go deeper
Know that a call can succeed while its response is lost, so a create that is simply retried can leave two items in the tracker for one local subject.
Explain why applying an update twice is harmless while creating twice is not, and describe searching for a marker the sync itself planted before it creates again.
Walk the durable-intent pattern end to end: a pending record written before the call, a token carried into the item, adoption on retry, and a per-subject guard against two workers racing.
Own the cleanup and the promise: how an existing tail of duplicates is reconciled without discarding anyone's work, and what your integration guarantees about one item per subject.
Every call across a boundary has three outcomes, not two: it succeeded, it failed, or it succeeded and you did not find out. The third is the interesting one, and creation is the operation where it hurts, because creation is the only common sync write that is not naturally repeatable. Applying a state update twice leaves the same state; creating twice leaves two items. ## The failure is the acknowledgement, not the write Reconstruct the incident precisely, because the wrong reconstruction leads to the wrong fix: 1. The sync decides an item must exist in the tracker for some local subject. 2. It sends the create. The tracker accepts it, allocates an identifier, and returns it. 3. The response is lost — a reset socket, a client-side timeout, a process killed mid-deploy, a gateway that gave up while the tracker was still working. 4. The sync has no identifier, so it stores nothing. 5. On the next pass, the local subject still has no recorded link, so the sync concludes the item does not exist and creates it again. Nothing there is a bug in the tracker, and no amount of retry tuning helps: a shorter timeout makes it likelier, a longer one makes it rarer but never impossible. The condition to remove is the silence in step 4 — the state where a create may have happened and nothing durable says so. ## Write the intent before you make the call The general fix is to stop treating "the link row exists" as the record of the create, and to keep a record of the **attempt** instead: - Before calling, generate a **token**, an opaque unique value derived locally, and persist a pending record: this subject, this token, attempted at this time. - Call the tracker, carrying the token somewhere the created item will keep it. - On a successful response, store the returned identifier and mark the record complete. - On any failure, including a timeout, leave the record pending. **Do not** retry blindly. A pending record converts "I do not know" into a durable question the next pass can answer. That pass no longer has the job "create"; it has the job "find out whether my earlier create landed". ## Making the create findable Answering that question needs the token to be visible from outside. Two routes: - **Native idempotency.** Some APIs accept a caller-supplied token on a create and, on seeing the same token again, return the original item rather than making a new one. This is the best answer available, because the deduplication happens on the server, inside the same transaction that created the item, with no window for the caller to lose. - **A searchable marker.** Where that is not offered, the sync writes the token into a durable, searchable part of the item it creates — typically inside the human-readable text it is already composing, in a form nobody would type by accident. The retry searches for the token first and adopts what it finds instead of creating. The marker route has a real race: two workers can each search, each find nothing, and each create. Serialize per subject — a lock, or a uniqueness constraint on the pending record's subject — so only one attempt is ever in flight for one local subject. | Approach | Deduplicates where | Race window | |---|---|---| | Retry the create with nothing carried | nowhere | every lost response | | Caller-supplied idempotency token | on the tracker | none for the caller | | Marker plus search-before-create | in your sync | between the search and the create | | Marker plus a per-subject lock | in your sync | only if the lock is lost | ## The duplicates you have already made A sync that ran without any of this leaves a tail of duplicates in the tracker, and they are not the tracker's problem to solve. Sweep for items carrying the sync's own marker, group them by local subject, and for each group keep the one with activity on it — a comment, a state change, an assignment — because that is the one people have been working in. Point the local subject at the survivor, and close the rest with a reference to it rather than deleting them, so anyone who bookmarked a duplicate lands somewhere that explains itself. ## The window that never fully closes Even with a token, one gap remains: the item can be created and the token stored, and the process can die before the local record is marked complete. That is fine, and it is the point. The next pass finds a pending record, searches, finds the item, and adopts it. The design goal is not to eliminate the crash window but to make every state it can leave behind **recoverable by re-running the same code**.
- Why is matching on the title not good enough to spot the item your earlier create already made?Titles are edited by people, truncated by whoever composed them, and legitimately duplicated across different subjects. A match may land on a ticket somebody else filed for an unrelated failure with similar wording, and adopting it attaches your local subject to work that has nothing to do with it. A token you generated is unique by construction and nobody edits it by accident.
- Where does this pattern apply beyond creating a tracker item?Anywhere the sync makes something with its own identity — a comment it posts on an existing item, or a second link row for the same pair. Each turns one lost response into a visible duplicate, and each is fixed the same way: durable intent written first, a token that travels with the thing created, and a retry that searches before it creates.
saying these in an interview costs you the question
- Assuming a timeout means the create never happened
- Retrying a create with no way to recognise it later
- Matching on titles to find the duplicate you made
- Deleting duplicate items instead of closing them
- Letting two workers create for the same subject at once