A suite runs as many parallel shards inside one pipeline run, and each shard pushes its own outcomes into a case repository. How would you design the cycle handle so the run lands as a single cycle?
answer
- N writers, one container
- who creates it, and when
- check-then-create is a race
- pass the handle as data
- shard index belongs in the key
basics
~20 sMint the cycle once in a setup job before the shards start, pass its handle down as an ordinary job value, and let shards record only. Get-or-create inside each shard races and yields duplicate cycles.
solid answer
~50 sSharding turns the cycle handle from a detail into a design decision, because N writers now need to agree on one container before any of them has anything to write. The robust shape is **mint-once-upstream**: a setup job creates the cycle, publishes the handle as an output, every shard reads it and records only, and a teardown job that runs regardless of shard outcomes performs the reconciliation and close. The tempting alternative — each shard derives a deterministic cycle name from the run identity and creates it if absent — races: two shards check at once, both find nothing, and you get two cycles for one run. It is only safe where the repository can create-if-absent atomically, which you must verify rather than assume. Whatever you choose, the shard index has to be part of each push's idempotency key, or a retried shard collides with its siblings.
go deeper
Know that parallel shards must all write into one container, and that the handle for it has to exist before any shard starts pushing rather than being invented by whichever shard finishes first.
Explain the race in check-then-create: two shards look at once, both find nothing, and both create. Say why passing a handle created upstream avoids it entirely.
Cover the operational edges — a teardown that runs even when shards fail, the shard index inside the idempotency key, and an explicit decision about what a cycle missing one shard should look like.
Own the framing that the cycle handle has an owner, a lifetime and a creation authority separate from the writers, and weigh the extra setup job against the coordination-free options that fail invisibly.
## Why sharding changes the problem With one runner, the cycle handle is trivial: whoever pushes creates or is handed the container. With N shards the same handle has to exist *before* any shard writes, be identical for all of them, and outlive shards that die. Three separate concerns fall out of that — who creates it, how it travels, and who decides the run is finished — and each has a wrong answer that looks reasonable on a whiteboard. ## Option one: mint once upstream A setup job runs before the fan-out, creates the cycle, and publishes its handle as an ordinary job output or environment value. Shards consume the handle and record only; they never create. A teardown job that runs whether the shards passed, failed or were cancelled does the reconciliation and any closing. - **Why it works.** Exactly one writer creates, so there is no race to lose. The handle is data flowing down the pipeline graph, which every CI platform already does well. - **What it costs.** An extra job on the critical path, and the setup job's identity needs authority to create a container that recording shards do not need — the permission split that separates recording from creating and closing is its own subject, but the design has to acknowledge it exists. - **The failure to plan for.** If the setup job succeeds and every shard then fails to start, you own an empty cycle that has to be cleaned up or explained. ## Option two: deterministic name with get-or-create Each shard derives the same cycle name from the run's identity and creates the cycle if it does not find one. - **Why it is tempting.** No extra job, no handle to pass, and each shard is self-sufficient — attractive when shards can be re-run individually. - **Why it races.** Two shards starting together both look, both find nothing, and both create. You end up with two cycles bearing the same intended name and half the results in each, which is worse than either failing outright because it looks like the run was reported. - **When it is legitimate.** Only where the repository can perform create-if-absent as a single atomic operation, or where you serialise creation behind your own ingest component. Verify this; do not assume a product provides it. ## Option three: a cycle per shard, rolled up later Each shard makes its own cycle and a reporting layer aggregates them. - It removes all coordination, which is genuinely attractive at large fan-out. - It forfeits the single-cycle view that made anyone want a case repository, multiplies the containers a human must navigate, and pushes the aggregation problem to a layer whose job is presentation, not truth. ## Comparing the three | Approach | Coordination needed | Race risk | Single-cycle view | |---|---|---|---| | Mint once upstream | A setup job and one passed value | None | Yes | | Deterministic get-or-create | None | High, unless creation is atomic | Usually, but not reliably | | One cycle per shard | None | None | No | ## The cross-cutting details that decide it in practice 1. **Retries must not collide.** A retried shard re-pushes into a cycle its siblings also wrote to, so the shard index belongs in the idempotency key alongside the run identity. Omit it and shard four's retry looks like shard one's original write. 2. **Partial completion needs a decision.** When one shard dies, the cycle holds most but not all of the run. Decide explicitly whether the teardown closes such a cycle, leaves it open, or marks it incomplete — and make the count gap between produced and accepted outcomes visible, because a short cycle is the failure nobody spots. 3. **Teardown must run unconditionally.** Reconciliation placed in a step that is skipped when shards fail is reconciliation that only ever runs on the happy path, which is where it is least needed. 4. **The handle should be data, not a convention.** A handle passed explicitly can be logged, echoed into the job summary, and reused by a replay job. A handle re-derived by each consumer from a naming rule silently changes meaning the day the rule changes. ## The judgment being tested The interesting part is not naming a mechanism; it is recognising that the cycle handle now has an owner, a lifetime, and a creation authority distinct from the shards that write into it. Mint-once-upstream states all three explicitly and pays one extra job for the privilege. The alternatives buy simplicity by leaving one of the three implicit, and each of them fails in a way that looks, from the outside, like the run reported fine.
- Every shard derives the same cycle name from the run identity and creates it if missing. What actually goes wrong?Two shards that start together both look, both find nothing, and both create — so one run produces two cycles with the intended name and the results split between them. Check-then-create is only safe when the repository makes it a single atomic operation, or when creation is serialised behind one component you control. Verify that property rather than assuming it.
- One shard dies without pushing. What should the teardown job do with the cycle?Decide deliberately and make it visible. Reconcile the outcomes the repository accepted against what the run should have produced, and mark or leave open a cycle known to be short rather than closing it as though it were complete. A closed cycle silently missing a shard's worth of results is indistinguishable from a smaller run that finished cleanly.
saying these in an interview costs you the question
- Lets every shard create the cycle if absent
- Assumes check-then-create is atomic
- Puts reconciliation in a step skipped on failure
- Omits the shard index from the idempotency key
- Treats a short cycle as an acceptable outcome