Which restart policy suits a long-running telemetry ingester, and which suits a run-to-completion import job?
answer
- after it ends, does another start?
- reacts to the ending, not the cause
- a service has no successful ending
- a job restarted on any ending loops
- never-restart trades recovery for a frozen instance
basics
~10 sA long-running service wants restarting whenever it ends, because any ending is abnormal. A run-to-completion job wants restarting only when it ends badly, so that a successful run is not repeated forever.
solid answer
~40 sA restart policy answers exactly one question: after this instance ends, does something start another? **Restart on any ending** creates a replacement however the instance finished — right for a long-running service, where even an orderly exit means the service has stopped serving. **Restart only on a bad ending** creates a replacement only when the attempt failed — right for a run-to-completion job, because a job that finished its work is done, and restarting on any ending would re-run it forever. **Never restart** leaves the ended instance alone, which is what you want when repeating the work unattended is not safe, or when you want the failed instance kept exactly as it is. What counts as a bad ending is the exit-status contract between the process and whatever supervises it.
code
yaml · 17 linesworkload:
name: telemetry-ingester
shape: long-running-service
restart:
when: any-exit # any-exit | failed-exit-only | never
initialDelay: 10s
growthFactor: 2
maxDelay: 5m
job:
name: nightly-import
shape: run-to-completion
restart:
when: failed-exit-only # any-exit here would re-run a finished import forever
initialDelay: 30s
growthFactor: 2
maxDelay: 10mgo deeper
Recall the three policies and the one question they answer: after this instance ends, does another attempt start? Match them to a service that should always be running and a job that should finish.
Explain why a job under restart-on-any-ending loops on success rather than on failure, and why an orderly ending from a long-running service is itself a symptom worth investigating.
Show the operating judgment: never-restart trades churn for a hard outage and gives up automatic recovery from transient causes, so it is a deliberate, alerted, time-boxed choice and not a way to quieten a workload.
Argue the default your platform should set for each workload shape, and what teams must decide for themselves — the cost of an unnecessary outage against the cost of a workload that silently repeats work nobody asked it to repeat.
## The policy is a reaction rule, not a repair A **restart policy** is a standing instruction attached to a workload, read by whatever supervises it at one moment only: when an instance ends. It cannot prevent anything. It decides whether another attempt begins, and that is the whole of its power. Every platform that runs containers has some form of it, because the alternative — a workload that quietly vanishes the first time a dependency blips — is not operable. The three policies in general use are the same three everywhere, whatever they are spelled: | policy | after an orderly ending | after a bad ending | the workload it fits | |---|---|---|---| | restart on any ending | yes | yes | a long-running service | | restart only on a bad ending | no | yes | a run-to-completion job | | never restart | no | no | work that must not repeat unattended, or an instance you want frozen | ## Why a long-running service takes the first row A telemetry ingester is supposed to be running now, and at every moment after now. Its useful states are "running" and "broken"; it has no successful ending. So when the instance ends, the reason barely matters — whether the process decided to stop, or was ended from outside, the service is no longer ingesting, and something should try again. Restarting on any ending encodes that: the desired state is *present*, and the policy keeps pushing toward it. This is also why a clean, orderly ending from a long-running service should make you suspicious rather than relieved. Something asked it to stop, or it decided its own work was over, and neither is normal for a workload whose job never finishes. ## Why a job takes the second row A nightly import has the opposite shape. It has a defined unit of work and a real successful ending, and "finished" is the outcome you wanted. Put it under restart-on-any-ending and you have built a loop out of a success: the job completes, the supervisor observes an ending, starts another attempt, that one completes too, and the import runs continuously until someone notices. This is one of the most common self-inflicted restart loops in production, and the giveaway is that nothing is failing — each attempt succeeds. The policy that fits reacts to the *outcome*, not merely to the ending: restart only when the attempt ended badly, so a failed import is retried and a successful one is left finished. ## When never-restart is the right choice Never-restart has two honest uses, and one dishonest one. - Honest: repeating the work unattended is not safe, so a failed attempt should wait for a person to decide whether re-running it is correct. - Honest: you want the failed instance preserved as it is, because the next attempt would displace the evidence you are about to read. - Dishonest: using it to silence a noisy workload. That converts an intermittently-available workload into an unavailable one and changes nothing about why it was failing. Never-restart also removes the automatic recovery that the other two policies provide for **transient** causes — a dependency that was briefly unreachable, a node that was briefly starved. Those are real, and they are the reason the default on most platforms leans toward restarting rather than not. ## What the policy interacts with Two interactions are worth naming, because they are what makes the choice non-obvious in practice: 1. **The delay between attempts.** A policy that restarts is almost always paired with a growing delay, so that a workload which cannot start does not attempt it thousands of times a minute. The policy decides *whether*; the delay decides *how often*. 2. **What forces the ending in the first place.** An instance can end because its process exited, or because a failing check asked for it to be replaced, or because the host reclaimed it. The policy reacts to the ending the same way regardless — which is precisely why the policy choice never tells you anything about the cause. ## A working rule Ask what the workload's *successful* state is. If success means still running, restart on any ending. If success means finished, restart only on a bad ending. If success means a human looked at it first, do not restart — and make sure something alerts, because a workload that is deliberately not restarting is also a workload that is deliberately down.
- Why does restarting on any ending turn a successful run-to-completion job into a loop?Because that policy reacts to the ending, not to the outcome. A finished import ends, the supervisor sees an ending and starts another attempt, and each attempt succeeds and ends again. Nothing is failing, which is why it is often noticed late.
- What do you give up by putting a long-running service on never-restart?Automatic recovery. Any transient cause — a dependency briefly unreachable, a host briefly starved — now becomes a full outage that lasts until a person acts. You gain a frozen instance to inspect and pay for it in availability.
- Does choosing restart-only-on-a-bad-ending make re-running the work safe?No. It only decides when another attempt begins. Whether repeating a partially-completed unit of work is safe is a property of the work itself, and if it is not, the policy will happily repeat it anyway.
saying these in an interview costs you the question
- Treats restart-on-any-ending as the safe default for every workload, including jobs
- Expects a completed job under restart-on-any-ending to stay completed
- Believes a restart policy prevents failures rather than reacting to them
- Assumes never-restart makes a failure visible on its own, with no alert
- Thinks restart-only-on-failure makes repeating the work automatically safe