skip to content

A workload never reached running and is serving nothing — which four distinct startup failures produce that symptom?

level: juniorimportance: must knowfreq 74%

answer

  1. four failures, one outside symptom
  2. artifact, start, death, readiness
  3. which step authored the message
  4. is there a termination record
  5. alive but held out of routing

basics

~20 s

Four separate failures look identical from outside: the image never arrived on the node, the start was refused before any code ran, the process started and died, or it runs and never reports ready. Naming which one is the first move.

solid answer

~50 s

Four things have to succeed before an instance serves, and a failure in any of them is summarised as "not running". First the artifact has to arrive on the node — an unauthenticated fetch, a rate limit, a reference resolving nowhere, or nothing matching the node's processor architecture stops it there. Second the runtime has to construct a running process — a value the spec references that does not exist, a mount it cannot attach, or a declared command that is not there stops it before any of your code runs. Third the process can start and die at once. Fourth it can run happily and never report ready, so nothing is routed to it. Each adjacent pair is separated by one cheap observation, so name the stage before forming any theory about the cause.

code

pseudocode · 22 lines
pseudocode
// classify the CURRENT state of one workload instance
lastStepReported  = status.lastStepThatReported   // "artifact fetch" | "start" | "process"
terminationRecord = status.lastTermination        // none, or a record
processAlive      = status.running
inRouting         = status.receivingTraffic

if lastStepReported == "artifact fetch" and not processAlive:
    return "1 - the artifact never arrived"

if lastStepReported == "start" and terminationRecord == none:
    return "2 - the start was refused before any code ran"

if terminationRecord != none and not processAlive:
    return "3 - it started and died"

if terminationRecord != none and processAlive:
    return "3 - it died at least once and is running again"

if processAlive and not inRouting:
    return "4 - running, but it never reported ready"

return "not a startup failure - alive and in routing"

go deeper

for a junior

Be able to say the four failures out loud in order — artifact, start, death, readiness — and give one observation that separates each pair. That structure alone is what the question is testing.

for a middle

Explain why evidence from a later stage is absent rather than negative at an earlier one, and why the author of the message you are reading is the fastest discriminator available.

for a senior

Show the triage order and what each step eliminates, and say what the evidence does not prove: retries appear at three of the four stages, and an empty stream narrows without settling.

for a principal

Frame it as a standard: agree the four names and the one observation per boundary so that anyone on call reaches the same stage from the same status, whichever platform the estate runs.

## Four stages, one symptom Between the moment a platform accepts your declaration and the moment an instance serves its first request, four things must succeed in order. A failure in any of them shows up in a summary view as some variant of *not running*, which is why the first move is to name the stage rather than to guess the cause. 1. **The artifact never arrived.** The node must obtain the image before anything can be created from it. That fetch can fail because it was never authenticated for that repository, because the registry applied a rate limit, because the reference resolves to nothing, or because nothing it found matches the node's processor architecture. No process of yours has run. 2. **The start was refused.** The node has the image and the runtime still cannot construct a running process: a value the spec references does not exist, a mount cannot be attached, or the declared command is not at the path given or is not executable. Again, none of your code has run. 3. **It started and died.** The process ran — for a second or for a minute — and ended. There is a termination record for that instance, and usually some output of its own. 4. **It runs and never reports ready.** The process is alive, its uptime is climbing, and the platform withholds traffic because the instance has never announced that it can serve. The four are ordered, and that ordering is the useful part: a failure at an earlier stage makes every later stage unreachable, so evidence that belongs to a later stage is *absent*, not negative. There is no output of yours to read at stages one and two because nothing of yours ran. ## The boundary between each adjacent pair Every boundary is settled by one cheap observation, and the cheapest of them all is **who authored the message you are reading** — the node agent's fetch step, the runtime's start step, or your own process. | Boundary | The question that settles it | It is the earlier stage when… | |---|---|---| | 1 against 2 | Which step reported the failure? | the report concerns obtaining the artifact, not starting it | | 2 against 3 | Is there a termination record for this instance? | there is none, and the only message came from the runtime | | 3 against 4 | Is a process alive right now? | it is not, and the instance is being replaced | | 4 against serving | Is it in routing? | it is alive, never replaced, and still receives nothing | ## The order to check, and why that order Run the checks in the order that costs least and rules out most, and stop at the first one that answers: 1. **Which step reported last?** A message about obtaining the artifact means you are at stage one, and nothing about your spec's values, mounts or command is implicated yet. 2. **Is there a termination record for this instance or its predecessor?** Its presence proves a process ran, which eliminates stages one and two outright. 3. **Is a process alive right now?** Alive with a climbing uptime rules out stage three for the current instance. 4. **Is it receiving traffic?** Alive and excluded is stage four, and it is the only one of the four where the instance itself is doing fine. Only after the stage is named does the cause list become short. Each stage has its own, largely disjoint, set of causes, and working the wrong list is how an hour disappears. ## Why the fourth failure fools people Stages one through three look like failures: something is red, something is being retried, something is missing. Stage four looks like success. The instance is alive, the output stream carries ordinary start-up lines, nothing is being replaced — and no traffic arrives. Engineers reading that picture often conclude the routing layer is broken, when the platform is behaving exactly as designed: it has no positive ready report, so it sends nothing. Anything waiting on that instance waits without an error to show for it. ## What this evidence does not prove - **Retries do not name a stage.** A failed fetch, a refused start and a process that dies are all retried; a climbing attempt counter is consistent with all three. What the retry policy is, and how it spaces attempts, is a separate subject. - **An empty output stream is evidence, not proof.** It favours stages one and two, but a process can also die before its first line is written. - **A stage is not a cause.** "The start was refused" is a diagnosis of *where*, not of *why*; it tells you which short list to read. - **One healthy node does not clear the workload.** An instance running somewhere else may simply have skipped a stage that another node cannot, which is its own diagnostic question. - **The absence of a termination record is not proof that nothing ran**, only the strongest cheap evidence available; platforms differ in how long they retain the record of a previous instance.

  • The attempt counter is climbing steadily — which of the four stages does that rule out?
    Only the fourth. A failed fetch, a refused start and a process that dies are all retried, so a climbing counter is consistent with the first three. An instance that runs and never reports ready is not retried at all: it stays alive and stays out of routing, which is why the counter sitting at zero is itself a clue.
  • The output stream for the current instance is completely empty — what does that narrow the diagnosis to?
    It favours the first two stages, where none of your code ran, over the third. It does not settle it: a process can die before writing its first line. Confirm with a termination record for this instance or its predecessor, which only exists if a process actually ran.
  • Why name the stage before reasoning about the cause at all?
    Because each stage carries a largely disjoint cause list, and all four present the same outside symptom. Naming the stage turns "anything could be wrong" into one short list — credentials and references, or values and mounts and the command, or the process's own early failure, or the readiness report.

saying these in an interview costs you the question

  • Calls every workload that will not start a crash loop
  • Debugs application code before any process has run
  • Treats a climbing attempt counter as proof of a crash
  • Assumes an empty output stream means the platform lost it
  • Never asks whether the artifact reached the node at all
  • Calls a running instance healthy while it gets no traffic