skip to content

An ingester has been in a restart loop for six hours; a teammate proposes never restarting it — what does that change?

level: seniorimportance: should knowfreq 40%

answer

  1. a symptom, not a diagnosis
  2. it changes the churn, not the cause
  3. trades intermittent availability for none
  4. loses automatic recovery from transient causes
  5. buys a frozen instance and steady evidence

basics

~20 s

It ends the churn and leaves the last failed instance in place to inspect, but it changes nothing about the cause and removes the automatic recovery that would have fixed a transient one. A restart loop is a symptom, not a diagnosis.

solid answer

~50 s

Switching the policy converts an intermittently-present workload into an absent one. That is a real trade and sometimes the right one — the churn stops, the last failed instance stays put instead of being displaced by the next attempt, and the evidence stops moving while you read it. But nothing about the failure changes, and if the cause is transient, you have just removed the mechanism that would have recovered the workload without you. The deeper point is what the loop itself establishes: only that the instance keeps ending and the policy keeps asking for another. It does not say whether the process ended itself or was ended from outside, why, or whether the trouble is inside the workload or in something it depends on. Writing "it is in a restart loop" as the cause in an incident summary is writing down the symptom.

go deeper

for a junior

Recall that a restarting workload tells you something is wrong but not what. The loop is the alarm, and the cause has to be found in the evidence the failing attempts left behind.

for a middle

Explain the two facts the loop composes — instances ending, and a policy asking for another attempt — and why several very different causes are indistinguishable at that altitude.

for a senior

Show the operating trade: freezing the workload preserves evidence and stops churn but converts intermittent availability into none, and it must be time-boxed, owned, announced and alerted rather than used to quieten a signal.

for a principal

Set the standard other teams follow: when containment justifies stopping a looping workload, who may decide it, and what an incident record must contain beyond the symptom that was easiest to see.

## What a restart loop actually establishes A restart loop is the visible composition of two facts, and only two: 1. Instances of this workload keep **ending** — by the process exiting on its own, or by being ended from outside. 2. The workload's **restart policy** asks for another attempt each time, so a new instance appears. That is the whole of the information content. It is a genuinely useful signal — it is usually the first thing anyone notices, and it reliably says *this workload is not working* — but it is a symptom that names no cause, and treating it as a cause is the classic mistake this question is asked to catch. ## What it does not establish Each of these is invisible from the loop itself: - **Which way it ended.** The process deciding to stop and something outside ending it produce the same loop from the outside. - **Why.** A configuration value, an unreachable dependency, a resource ceiling, a failing check that forces replacement — all of them look identical at this altitude. - **Where the trouble is.** A workload failing because of its own defect and one failing because the thing it connects to is down are the same loop. - **Whether it ever ran at all.** An attempt that never reached the workload's own code and one that ran for a minute both end; distinguishing them is a separate subject with its own evidence. - **Whether every attempt fails the same way.** Later attempts often fail for reasons the earlier ones created — an unreleased connection, a half-written file, a claim not given back. So the loop is where a diagnosis starts, never where one ends. ## What the never-restart proposal buys and costs | | keep restarting | stop restarting | |---|---|---| | churn on the host and its dependencies | continues, at the capped rate | ends | | the last failed instance | displaced by the next attempt | stays, available to inspect | | recovery if the cause is transient | automatic, at the next attempt | none until a person acts | | availability | intermittent, brief windows | none at all | | the cause | unchanged | unchanged | Read the last row first. Both columns are identical there, and that is the answer to the question: the proposal is a choice about **churn against availability**, not a step toward a fix. It earns its place in two situations. One is evidence preservation — if each new attempt is displacing the output you need, freezing the workload stops the record moving. The other is blast radius: a workload whose every attempt hammers a struggling dependency, or claims and releases a shared resource, can be making a wider incident worse, and stopping it is a containment action taken deliberately. In both cases it is a time-boxed decision with an owner, announced to whoever depends on the workload, and paired with an alert — because a workload that is deliberately not restarting is a workload that is deliberately down, and nothing will page anyone about it. ## Where it is the wrong move When it is used to make a noisy signal stop. A loop is loud precisely in proportion to how broken the workload is; silencing it removes the alarm and leaves the fire. The tell is the phrasing: "turn off the restarts so the alerts stop" is quieting an alarm, while "freeze the current instance so I can read it, back in ten minutes" is a diagnostic step with a stated end. ## A defensible order of moves 1. **Take the evidence first**, while the workload is still cycling: the previous attempt's retained output, and the current delay between attempts as a measure of how long this has been going on. 2. **Decide whether the loop is harming anything else.** If each attempt is loading something already struggling, containment is a real argument for stopping it. 3. **Only then change the policy**, and only with an owner, a time box and an alert — and say out loud that this converts intermittent availability into none. 4. **Write the cause, not the symptom.** "The ingester was in a restart loop" is what you saw. What you owe the incident record is what ended the instances, and why. ## The judgment being tested The question is not really about the policy field. It is about whether you can separate a loud symptom from a cause under time pressure, and whether you will trade a system's automatic recovery away without noticing that you did. A candidate who accepts the proposal without naming what it costs, or who rejects it without conceding what it buys, has missed a different half of the same point.

  • What does a restart loop tell you that is genuinely useful?
    That instances keep ending and the policy keeps replacing them — so the workload is not working, and it has been failing repeatedly rather than once. The current delay between attempts also summarises how long the loop has run without a stable instance.
  • When is stopping the restarts a legitimate diagnostic step rather than silencing an alarm?
    When each new attempt is displacing the evidence you are reading, or when the attempts are loading a dependency that is already struggling. It stays legitimate when it is time-boxed, owned, announced, and paired with an alert.
  • Why can the tenth attempt in a loop fail differently from the first?
    Because earlier attempts leave state behind — an unreleased connection, a partially written file, a claim not given back. The later failures can be consequences of the loop itself rather than of whatever started it.

saying these in an interview costs you the question

  • Records 'it is in a restart loop' as the cause in an incident summary
  • Assumes a restart loop always means the workload's own code is broken
  • Treats switching restarts off as a fix rather than a trade of churn for downtime
  • Assumes a looping workload is harmless to everything around it
  • Reasons only from the newest attempt, as if every attempt failed identically
  • Stops the restarts without telling anyone who depends on the workload