skip to content

When two versions of a job's logic run side by side over the same input and only one publishes, which effects of the silent arm still reach the outside world?

level: seniorimportance: should knowfreq 47%

answer

  1. publishing is a thing you prevent
  2. not just the output path
  3. the reader's recorded position
  4. side effects the step graph cannot show

basics

~20 s

Everything the silent arm does besides writing its own output: a shared destination, an advanced reader position on the input, side effects fired from inside the logic, telemetry keyed by job name, and doubled load on shared capacity.

solid answer

~50 s

Running two versions side by side — both over the same input, only one publishing — makes *not publishing* something you have to engineer, not something you configure once. The obvious path is the output destination, and it is the one everybody remembers. The ones that bite are the others: if both arms share the identity that records how far the input has been consumed, the silent arm can advance it and the publishing arm skips records. An **opaque step** — a hand-written per-record function the engine can only call, never inspect — that posts to a service or sends a message will do it twice. A consumer watching for new output will react to whatever lands anywhere it watches. The silent arm's failures page the on-call unless every signal says which arm produced it. And two arms cost roughly double the compute and read the source twice.

go deeper

for a junior

Recall that the second version must not write where readers look, and must not share the identity that records how far the input has been read. Those two mistakes account for most of the damage.

for a middle

Explain why a hand-written per-record function is the dangerous case: the engine can only call it, so its external effects never appear in the job's step graph and happen once per arm.

for a senior

Demonstrate that you enumerate effects before starting: destination, reader identity, side-effecting steps, watched locations, telemetry, shared capacity. Then explain warm-up as a distortion to exclude rather than a difference to investigate.

for a principal

The call is how much isolation to pay for. Fully isolated infrastructure removes the risk and also removes the realism; say where your organisation draws that line and who is accountable when a silent arm turns out not to be silent.

## Not publishing is something you build A side-by-side run means both versions of the transformation logic execute over the same input at the same time, with only one of them reaching readers. The phrase *only one publishes* describes an outcome, not a switch. Publishing is not a single act on most platforms: it is a write to a destination, plus whatever the write causes, plus whatever the logic does on its own along the way. Each of those is a separate thing to contain, and the ones that get missed are never the destination. What counts as publishing also varies by platform: committing a set of files to shared storage, updating rows in a store that something queries, or emitting records that a downstream reader consumes. Contain all three shapes you actually have, not the one you happen to be picturing. ## The paths by which the silent arm still reaches the world 1. **The destination.** The obvious one. Give the candidate arm its own output location and make sure nothing defaults the writer back to the real one. It must, however, keep the same write *mechanism* and the same output shape, or you end up comparing writers instead of logic. 2. **The input reader's recorded position.** If both arms read through a shared identity that records how far the input has been consumed, the silent arm can mark records as read and the publishing arm will never see them. Give each arm its own reader identity, or feed the candidate from a stored copy of the same records. The mechanics of how a log or queue records consumer progress are a messaging subject, not this one; what matters here is that the identity must not be shared. 3. **Side effects inside the logic.** An **opaque step** is a hand-written per-record function the engine can only call, never inspect or rewrite. If one of them posts to an external service, sends a message, or writes an audit row, it fires twice — once per arm — and nothing in the job's step graph shows it, because the engine cannot see inside the function. Route these to a discard target in the candidate arm, or make the arm identity an input the function honours. 4. **Downstream triggers.** A consumer that reacts to a new file appearing, or to a row changing, will react to the candidate arm's output if it lands anywhere watched. Being in a different folder under the same watched prefix is not isolation. 5. **Telemetry and alerting.** Failures, lag and throughput from the candidate arm land on the same dashboards and page the same on-call unless every signal carries which arm emitted it. A noisy candidate arm that wakes people at night gets switched off before it has told you anything. 6. **Shared capacity and the source.** Two arms roughly double the compute and read the source twice. On a shared pool, the candidate can starve the publishing arm; on a rate-limited source, it can throttle it. The publishing arm keeps priority, always. | leak path | what you see | containment | |---|---|---| | shared destination | readers get the candidate's numbers | separate location, no default back to the real one | | shared reader identity | published output silently misses records | separate identity, or a stored copy of the input | | side-effecting per-record function | messages and rows written twice | discard target, or arm identity passed into the function | | watched output location | downstream jobs run on the candidate's output | write outside every watched prefix | | untagged telemetry | the on-call is paged for a job nobody depends on | tag every signal with the arm, route candidate alerts elsewhere | | shared pool or rate-limited source | the publishing arm slows or throttles | priority to the publishing arm, cap the candidate's width | ## The distortion that is not a leak: warm-up A long-lived candidate arm starts with nothing in its **retained set** — everything a running job holds between records, such as counters, buffered join sides and the last value per key. Its earliest outputs are therefore different from the publishing arm's for a reason that has nothing to do with the logic change, and reading those rows as a real difference wastes a day. This varies sharply by runtime model: where continuous work runs as a rapid succession of small finite runs, each run is bounded and the distortion is short; where the runtime is record-at-a-time with a key-bound retained set, the arm may take a long time to become comparable; and a periodic job over a finite input has no warm-up at all. Either discard the warm-up window from the comparison or start the candidate arm early enough that it is settled before the window you intend to read. ## What must stay identical The candidate arm should differ from the publishing arm in exactly three respects and no others: - **its destination**, which must sit outside every watched location - **its reader identity**, so it cannot consume on the publishing arm's behalf - **its telemetry tag**, so its failures and lag are attributable and routed away from the on-call Everything else stays the same — same engine, same input, same period, same write mechanism, same output shape, and ideally the same width. Every further difference is a difference the comparison will faithfully report and you will then have to explain, and each one you allow costs you a day of reading rows that were never about the change.

  • The candidate arm's first two hours disagree wildly with the publishing arm, then converge. What is the likeliest explanation?
    Warm-up rather than logic. A long-lived arm begins with an empty retained set and with groups already half over, so its early output is incomplete by construction. Discard the warm-up window and compare only periods the candidate arm saw from beginning to end.
  • What must be true of the candidate arm's write path for the comparison to mean anything?
    It must use the same write mechanism and produce the same output shape as the publishing arm, differing only in destination. If the candidate writes a different format, or a different grain, differences may come from the writer rather than from the logic, and every row of the comparison becomes unattributable.
  • Is it acceptable for the candidate arm to run at lower parallelism to save money?
    For values, usually yes: a correct job should produce the same result at any width. But it changes the number of output pieces, the timing, and any behaviour that depends on how records are grouped across workers — so if the change touches ordering, ties or anything width-sensitive, match the widths or you will be explaining noise.

saying these in an interview costs you the question

  • Assumes pointing the output elsewhere makes the second arm harmless
  • Shares the input reader's recorded position between both arms
  • Forgets that a hand-written per-record function can call external services
  • Treats the second arm's warm-up period as a real difference in logic
  • Leaves the second arm's failures wired to the on-call alert
  • Ignores that two arms compete for the same cluster and the same source