skip to content

The same nightly batch reaches running on one node and never starts on another — what does that asymmetry tell you?

level: seniorimportance: should knowfreq 42%

answer

  1. same spec, different node
  2. a node's contribution is small
  3. a held image skips the fetch
  4. credentials, limits, mounts, architecture
  5. force the good node to fetch

basics

~20 s

Asymmetry points at node-local state rather than the spec: an artifact already held on one node, fetch credentials or a rate limit that differ per node, a mount attachable only in one failure domain, or nodes of differing processor architecture.

solid answer

~50 s

The declared workload is identical on both nodes, so the variable is the node, and there are only a few things that genuinely differ. A node that already holds the image never performs the fetch, so a credential or rate-limit failure cannot surface there. The credential a node presents when it does fetch can differ. A mount may be attachable only within one failure domain. Nodes can differ in processor architecture, so the image that resolves on one resolves to nothing usable on the other. Each of those maps to one of the four startup stages, so the move is to compare the two nodes on exactly those points rather than re-reading a spec that is provably the same. And do not over-read one success: the working node may simply be hiding a fetch failure behind a copy it already has.

go deeper

for a junior

Notice that the same declaration went to both nodes, so the difference has to be in the node. Ask what the failing node had to do that the working one did not.

for a middle

Enumerate what actually differs between nodes — a held image, credentials, attachable storage, architecture — and map each to the startup stage it would break.

for a senior

Show the cached-image trap and the falsifiable test: predict the stage, check both nodes, force the working node into the same position, and repeat once because a rate limit is time-dependent.

for a principal

Treat the asymmetry as a fleet property rather than an incident: a failure that only appears on fresh nodes surfaces during growth or a rebuild, so it belongs in what the estate verifies before it needs the capacity.

## Why identical declarations produce different outcomes The workload spec is one document and both nodes received it, so nothing in the document can explain a difference between them. Something in the node's own state has to. That narrows the search dramatically, because a node's contribution to a start is small and enumerable — what it already holds, what credentials it can present, what it can attach, and what it can execute. The discipline is to stop reading the spec, list what differs between the two nodes, and map each difference onto the stage it would break. ## What genuinely differs between two nodes | Node-local difference | Stage it breaks | Evidence on the failing node | |---|---|---| | It already holds the image; the other must fetch it | 1, the artifact never arrives | the failing node's last message comes from the fetch step | | The credential it presents to the registry | 1 | an unauthenticated or rejected fetch, only there | | Its share of a rate limit the registry applies | 1 | the failure moves around and comes back; it is time-dependent | | Its processor architecture | 1 or 2 | nothing usable resolves for that node, or the node cannot execute the file it was given | | What it can attach — a store reachable only in one failure domain, or a host path that exists on one machine | 2, the start is refused | a start-step message naming the mount | | What the process finds locally once running | 3 or 4 | the process runs on both, and fails or stalls on one | The last row is the reminder that asymmetry does not always mean an early stage. A process can start on both nodes and die on one, or run on both and become ready on only one, if the thing it depends on is reachable from only one of them. ## The cached-artifact trap This is the asymmetry that misleads most often, and it is worth stating on its own. A node that already holds the image does not fetch it, so every failure mode of fetching is invisible there. The workload runs, and the obvious conclusion — "the spec is fine, it runs over here" — is exactly wrong. What you have proved is that the image *can* run, not that it can be obtained. The consequences are specific and unpleasant: - Every node added later fails, because every new node must fetch. - Any replacement that lands somewhere new fails, while replacements that land where the image is already held succeed, which looks like randomness. - The failure appears at the worst moment, when capacity is growing or a node has been rebuilt. The tell is that the set of working nodes correlates with age or with prior use of that image, rather than with anything in the workload. ## Turning the asymmetry into a test A hypothesis that explains the difference is not yet a finding. Make it falsifiable: 1. **Name the difference and the stage it predicts.** "The failing node must fetch, so I expect its last message to come from the fetch step, and the working node's not to." 2. **Check the prediction on both nodes**, not only the failing one. Half the evidence lives on the node that works. 3. **Force the working node into the same position.** Have it start something that references an artifact it does not already hold. If it now fails the same way, the node was never special — the cached copy was. 4. **Repeat once.** A rate limit, in particular, is time-dependent: a node that failed a minute ago can succeed now, and a single observation will happily confirm a theory that is wrong. ## What the asymmetry does not tell you - **It does not clear the spec.** A value referenced but missing everywhere would fail everywhere; a value present on one node's reachable store and not another's fails asymmetrically for a node-local reason. - **It does not always mean a node-local cause.** Two attempts separated in time can differ because a limit was in force for one of them, so reproduce before concluding. - **It does not name the stage by itself.** It narrows the cause list to node-local things; you still name the stage from the same four-way evidence you would use anywhere. - **It does not mean the working node is correctly configured.** It may simply not be exercising the step that fails — which, for a fetch, is the usual case.

  • The workload starts on every node that already holds the image and fails on every fresh one — which stage is it?
    The first: obtaining the artifact. The nodes that work never perform the fetch, so they cannot exhibit its failure. Expect the cause to be a credential, a rate limit, or a reference that resolves nowhere, and expect every node added from now on to fail the same way.
  • How do you keep a node-local hypothesis honest?
    By predicting before you look and by testing both directions. Say which stage the difference should break, confirm it on the failing node, confirm its absence on the working one, then force the working node into the same position — for a cached image, by having it start something it does not already hold.
  • Which node-local differences can break a stage after the process is already running?
    Anything the process itself reaches for. A store, a peer or a credential source available from one failure domain and not another lets the process start on both nodes and then die on one, or run on both and become ready on only one. Asymmetry is not proof of an early stage.

saying these in an interview costs you the question

  • Concludes the spec is correct because one node runs it
  • Ignores an already-held image when comparing two nodes
  • Assumes every node presents the same fetch credential
  • Assumes every node shares one processor architecture
  • Draws a conclusion from a single timed-out attempt
  • Reads the failing node only, never the working one