skip to content

A CI job that pushes an automated run's outcomes into a case repository times out and is retried, and the cycle ends up holding the same outcomes twice. What mechanism prevents that, and how do you choose its key?

level: seniorimportance: must knowfreq 58%

answer

  1. a timeout is not a failure
  2. at-least-once delivery, once-only recording
  3. the key names the write, not the attempt
  4. stable on retry, distinct per run
  5. one key per chunk, not per push

basics

~20 s

Send an idempotency key with the push, derived from the run's identity rather than generated fresh on each attempt, so a repeated write is recognised as the same one. Timeouts are ambiguous: the first call may already have landed.

solid answer

~50 s

The double-record comes from an ambiguous failure, not a buggy one. A timeout tells you the response was lost, never that the write was — so the retry pushes outcomes that may already be recorded. The fix is an **idempotency key**: a value derived deterministically from the run's identity (pipeline run identifier, plus shard index where the suite is sharded) and sent with the push, so the receiving side can recognise a repeat and record it once. The critical property is the derivation, not the format: the key must be *stable across retries of the same attempt* and *different for a genuinely new run*, which rules out anything random, anything timestamped at call time, and anything that includes the attempt counter. Where the repository offers no dedupe of its own, the equivalent is to reconcile — read what the cycle already holds and write only the difference — or to scope each pipeline attempt to its own cycle.

code

bash · 9 lines
bash
PIPELINE_RUN_ID="${PIPELINE_RUN_ID:?run identity required}"
SHARD_INDEX="${SHARD_INDEX:-0}"
CHUNK_SEQ="${CHUNK_SEQ:-0}"

# Stable across retries of the same work, distinct for a genuinely new run.
# Never fold in an attempt counter, a timestamp, or a random value.
IDEMPOTENCY_KEY="$(printf '%s:%s:%s' "$PIPELINE_RUN_ID" "$SHARD_INDEX" "$CHUNK_SEQ")"

echo "$IDEMPOTENCY_KEY"

go deeper

for a junior

Know that a retried job can record the same outcomes twice, and that the guard is a caller-supplied key sent with the push so the receiver can tell a repeat from a new write.

for a middle

Explain why a timeout is ambiguous rather than negative, and state the two properties a key must have: unchanged when the same work is retried, different when the work is genuinely new.

for a senior

Show the derivation you would actually use, including per-chunk keys for a chunked push, and give the fallback when the product offers no dedupe — read-and-reconcile, or a cycle per pipeline attempt.

for a principal

Own the delivery-guarantee argument end to end: the channel is at-least-once, so effectively-once recording has to be designed for, and decide whether that guarantee lives in the vendor, in your ingest step, or in the container model.

## Why a retry duplicates in the first place A failed push has three possible truths behind it, and the client can only distinguish two of them: 1. **The request never arrived.** Nothing was written. Retrying is correct and necessary. 2. **The request arrived and was rejected.** Nothing was written, and the response said so. Retrying is pointless until the cause is fixed. 3. **The request arrived, was written, and the response was lost.** The write happened. Retrying writes it again. A connection timeout, a dropped socket, a proxy cutting a long bulk call, a runner evicted between send and receive — all of these look identical from the client and all of them are case 1 or case 3. Because a CI job that fails is retried, either by an operator clicking the button or by an automatic policy, case 3 is not a rare edge; it is a routine event on any pipeline that pushes results. The visible symptom is a cycle holding two rows for the same item in the same round, which corrupts every count computed over that cycle. ## The mechanism An **idempotency key** is a caller-supplied value that identifies *the write*, not the attempt. The receiving side keeps a record of the keys it has already applied and, on seeing one again, returns the original result instead of performing the work a second time. That converts an at-least-once delivery channel into effectively-once recording, which is the strongest guarantee available when the network can lose responses. The key sits alongside the payload's three handles — the cycle, the item, the outcome — and answers a fourth question: *is this the same write I already accepted?* ## Choosing the key The format hardly matters; the derivation is everything. A workable key satisfies both halves of one rule: - **Stable across retries of the same work.** Re-running the failed job must produce the same key, or the mechanism does nothing. - **Distinct for genuinely different work.** A deliberate re-execution of the suite must produce a different key, or the second run's outcomes are silently swallowed as duplicates — a much worse failure, because it looks like everything worked. | Key derived from | Stable on retry? | Distinct per real run? | Verdict | |---|---|---|---| | A random value per call | No | Yes | Useless — every attempt is a new write | | A timestamp taken at call time | No | Yes | Same defect, harder to spot | | Pipeline run identifier plus shard index | Yes | Yes | The workable choice | | Pipeline run identifier plus attempt number | No | Yes | Retries look like new writes | | The commit SHA alone | Yes | No | Two deliberate runs of one commit collapse into one | | A hash of the payload contents | Yes | Mostly | Works until a message or duration differs by a byte | The practical answer in almost every pipeline is a stable run identity the platform already provides, combined with whatever partitions the work (shard index, job name, chunk sequence number when a bulk push is chunked). Each chunk needs its own key, or a retry of chunk five will be mistaken for chunk one. ## When the repository does not offer the mechanism Not every product exposes idempotent result recording, so you need fallbacks that achieve the same outcome with the guarantees you do have: - **Reconcile instead of append.** Read what the cycle already holds for the items in this batch, and push only the difference. This costs an extra read and races if two writers push at once, but it is simple and needs nothing from the vendor. - **Make the write itself replace rather than add.** Where a repository records at most one outcome per item per cycle, a repeat is naturally harmless; verify this rather than assuming it, since many products deliberately append every submission as history. - **Give each pipeline attempt its own container.** If a retry writes into a fresh cycle, a duplicate is impossible by construction. The cost is a proliferation of near-identical cycles and a reporting layer that must know which one is authoritative. - **Put the dedupe in your own ingest step.** A small wrapper that owns the key table gives you the guarantee regardless of what the product supports, at the cost of a component to run. ## The habit to keep Treat every push as retryable and every timeout as *possibly written*. Never respond to an ambiguous failure by skipping the retry to be safe — that trades duplicate rows for missing ones, which is the same corruption pointing the other way. Retry with the same key, and let the key decide. Where a suite is retried in whole or in part, the placement of that retry and what a pass on a second attempt conceals are separate concerns with their own answers; what matters here is only that repeated delivery of the same outcomes must record once.

  • Someone folds the job attempt number into the key so each attempt is traceable. What breaks?
    Every retry then produces a new key, so the receiving side sees a new write and records the outcomes again — the mechanism is present but inert, which is worse than absent because the team believes it is protected. Traceability belongs in metadata attached to the push or in the log, never in the value that decides whether two writes are the same.
  • You cannot tell whether an ambiguous push landed. Why not just skip the retry?
    Because skipping trades duplicated rows for missing ones, and a short cycle is harder to notice than a doubled one — nobody counts a cycle that merely looks small. With a stable key the ambiguity stops mattering: retry unconditionally and let the key collapse the repeat. Without one, read the cycle back and push only the difference.

An idempotency key is the serial number on a cheque. Present the same cheque twice and the teller sees the serial and moves the money once; hand over a freshly written cheque instead and it is simply a second payment.

saying these in an interview costs you the question

  • Generates the key fresh on each attempt
  • Treats a timeout as proof nothing was written
  • Skips the retry rather than deduplicating
  • Uses one key for a whole chunked push
  • Keys on the commit alone, collapsing deliberate reruns