skip to content

Lifecycle & Termination

How a container starts, is judged healthy, is asked to stop and is killed if it does not. Most production surprises here are about ending badly, not about starting.

on this pageshow

questions

21

A telemetry ingester is in a restart loop, but its logs come back nearly empty — where is the failing attempt's output?

level: juniorimportance: must knowfreq 72%

answer

  1. fresh instance, fresh stream
  2. you are reading the wrong attempt
  3. the failure ended before you looked
  4. ask for the previous attempt's output
  5. usually one prior attempt is kept

basics

~20 s

Each restart attempt is a new instance with a fresh output stream, so a request for the logs returns the attempt that just started, not the one that failed. Read the previous attempt's retained output instead.

solid answer

~40 s

A restart does not continue the old instance; it creates a new one, and that new instance's standard output and error start empty. By the time you look, the current attempt may be seconds old and still starting, so its stream holds a start-up banner and nothing else — the error you want was printed by the attempt that already ended. Platforms keep the immediately previous instance's stream for exactly this reason and give you a way to ask for it, so the first move on any looping workload is to request the *previous* attempt's output rather than the current one. Retention normally covers only that one prior attempt, so a loop steadily overwrites its own evidence; anything you need to keep has to be copied off the host while it still exists.

go deeper

for a junior

Recall that every restart attempt is a brand-new instance with an empty output stream, and that the failing attempt's output is retained separately and has to be asked for by name.

for a middle

Explain why the current stream is nearly empty at any moment during a loop, and note that retention usually covers only the immediately previous attempt, so a loop overwrites its own evidence.

for a senior

Show how you preserve the evidence before it is displaced: take the copy at first contact, read how the stream ends rather than only what it says, and distinguish the original failure from secondary ones later attempts introduce.

for a principal

Frame it as an observability requirement rather than a debugging trick. A workload whose only failure evidence dies with the instance is undebuggable at fleet scale, whatever the underlying cause turns out to be.

## What a restart actually creates A restart loop is a sequence of separate attempts, not one instance having a long bad day. When the first process inside a container ends — on its own, or because it was ended from outside — that instance is over. Whatever supervises the workload consults its **restart policy**, and if the policy asks for another attempt, a **new instance** is created from the same image and the same spec. That new instance has its own process and, the part that matters here, its own **standard output and error**: the two streams that are the log contract for a container on every platform. Nothing is carried across that boundary. The new attempt's streams begin empty and fill from its first line of output onward. ## Why the current stream is nearly always empty Now look at the timing. The thing you want to read happened in the attempt that has **already ended**. The attempt you can address right now is the one that has **just begun**. If the ingester accepts data for forty seconds and then dies, and you ask for its logs at a random moment, you are on average twenty seconds into a fresh attempt that has printed a banner and a configuration line and has not reached its failure yet. Two conclusions are available, and people routinely draw the wrong one: - The right conclusion is **you are reading the wrong attempt**. - The wrong conclusion is **this workload does not log anything** — a belief people then act on by adding logging to a workload that was already telling them what was wrong, in a stream they never opened. ## Where the evidence lives | where | what it holds | how long it lasts | |---|---|---| | the current attempt's stream | the attempt that has not failed yet | until this attempt ends | | the retained previous-attempt stream | the failure you are actually looking for | until the next attempt replaces it | | a copy taken off the host | each attempt that was copied in time | as long as the destination keeps it | Platforms expose the middle row deliberately — the retained output of the **immediately previous** instance — because this failure mode is universal and nobody can debug a loop without it. The names and the mechanisms differ between platforms; the idea does not. The important limit is in the third column. Retention normally covers **one** prior attempt. A restart loop is therefore self-erasing: attempt 41 overwrites the record of attempt 40, and by the time a human is looking, the *first* failure is long gone. That first failure is often the most informative one, because later attempts can be failing for a secondary reason — a half-written file, a connection the previous attempt never released, a queue position already claimed — that has nothing to do with what started the loop. ## The order to read in 1. Ask for the **previous attempt's** retained output, not the current attempt's. 2. Read it to the end. The last lines written are the ones closest to the ending, and they are the ones that matter. 3. Notice *how* the stream ends. A message about shutting down and an abrupt stop mid-line mean very different things about how the attempt finished. 4. If the previous attempt's stream is genuinely uninformative, stop re-reading streams and change the experiment: arrange for the output to be captured somewhere durable before the next attempt overwrites it. ## What this does not tell you Being disciplined about *which* attempt you are reading is a precondition for a diagnosis, not a diagnosis. The retained stream tells you what the process printed. It does not tell you whether the process chose to end or was ended from outside, and it will be empty rather than misleading if the attempt never got as far as running the workload's own code at all. Those are separate subjects with their own answers. What this leaf owns is the discipline itself: **in a restart loop, the currently running instance is the least interesting one on the host.** One practical note on the shape of the problem. The delay between attempts usually grows as the loop continues, and that delay is also what gives you time to read. Early in a loop, attempts arrive seconds apart and the retained stream is replaced while you are still scrolling it. Later, when the delay has grown to minutes, the same request returns a stable answer. That is not a reason to wait for the loop to slow down — it is a reason to take a copy the first time you have one, because the comfortable reading window opens exactly when the evidence is at its oldest and least representative.

  • Why does the current attempt of a looping ingester usually show only a start-up banner?
    Because that attempt is seconds old. Its streams began empty when the instance was created, and it has not yet run long enough to reach the point where the previous attempts failed. The banner is simply the first thing every attempt prints.
  • You want the output of the attempt before the previous one. Why is it usually unavailable?
    Host-side retention normally keeps one prior attempt, so each new instance's stream displaces the record two steps back. Anything older survives only if a copy was taken off the host before it was displaced.
  • The previous attempt's retained stream is empty too. What does that narrow down?
    It says the attempt produced no output before it ended — so either it never reached the workload's own code, or it was ended from outside before it printed anything. It does not by itself say which; both need evidence from outside the stream.

saying these in an interview costs you the question

  • Assumes the running instance's stream contains a crash that already happened
  • Concludes the workload logs nothing because the current attempt's output is empty
  • Thinks a restart resumes the same instance and continues its output stream
  • Expects every earlier attempt in the loop to still be retrievable from the host
  • Tries to open an interactive session inside an instance that lives for seconds
open as a page

When a running instance fails a liveness check, versus when it fails a readiness check, what does each failure cost?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Failing a liveness check costs the instance its life: the platform ends the process and starts a fresh one, discarding everything in memory. Failing a readiness check costs it only traffic: it leaves the routing set, keeps running, and rejoins when it passes again.

open as a page

When a nightly batch container stops, what does its exit code tell whatever supervises it?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The exit code is the container's own verdict on its run: zero means the work completed, any non-zero value means it did not. At stop time the supervisor branches on that one number, not on anything the job printed.

open as a page

Which restart policy suits a long-running telemetry ingester, and which suits a run-to-completion import job?

level: middleimportance: must knowfreq 56%

basics

~10 s

A long-running service wants restarting whenever it ends, because any ending is abnormal. A run-to-completion job wants restarting only when it ends badly, so that a successful run is not repeated forever.

open as a page

When a replica is asked to stop, why does closing its listening socket first still drop requests, and what order avoids that?

level: middleimportance: must knowfreq 60%

basics

~20 s

Because removal from the routing set is not instant. For a few seconds after the stop request, callers are still being sent here, and a closed listener turns each one into a connection failure. Leave the routing set first, keep serving while that propagates, then refuse new work.

open as a page

A ledger write replica is asked to stop, then forcibly killed 30 seconds later - what is that grace period for, and what happens when it expires?

level: middleimportance: must knowfreq 68%

basics

~20 s

The grace period is a budget for finishing work already in progress. The workload is asked to stop but keeps running; when the window ends it is killed outright. Anything still running at that instant is cut where it stood.

open as a page

A search indexer that needs four minutes to load its index is killed at the same point on every attempt — why?

level: middleimportance: must knowfreq 62%

basics

~20 s

The liveness check starts counting before the index is loaded, so the same failure count is reached at the same elapsed time on every attempt and the platform restarts it — forever. The fix is a start-up allowance, not a looser restart threshold.

open as a page

A container's first process is a wrapper script that launches a report renderer as a child; why is the renderer force-killed on every stop?

level: middleimportance: must knowfreq 64%

basics

~20 s

A stop request is delivered only to the container's first process, which here is the wrapper. The wrapper never passes it to the renderer, so the renderer works on unaware until the grace window expires and the forced kill ends it.

open as a page

A container is created, runs, stops and is later removed — what triggers each of those transitions?

level: middleimportance: must knowfreq 60%

basics

~20 s

Creation builds the container from its spec without running anything; a start request executes its first process, which is what running means; the container stops when that process ends, by itself or because it was asked to; removal is a separate request that destroys the record.

open as a page

Why does the delay between restart attempts grow, and what makes that growing delay reset?

level: middleimportance: should knowfreq 48%

basics

~20 s

The delay grows so that a workload which cannot stay up stops consuming the host and its dependencies at full speed. It resets only once an instance has stayed up long enough to count as stable, not merely because it started.

open as a page

On a liveness check, how do the failure threshold, interval and per-attempt timeout trade detection speed against false restarts?

level: middleimportance: should knowfreq 45%

basics

~20 s

Detection takes roughly the interval times the failure threshold, plus one timeout. Shrinking any of the three catches a hang sooner and also turns an ordinary stall into a restart, so all three should be set against the longest pause the service legitimately has.

open as a page

A long-lived container's process list fills over days with finished helper entries until new ones fail — why?

level: middleimportance: should knowfreq 42%

basics

~20 s

Nothing is collecting the exit statuses of the children that end inside the container. Each finished child keeps a process-table entry until its status is read, and with the first process not doing that job the entries accumulate until the table is exhausted.

open as a page

When a stopped container is restarted in place rather than replaced with a fresh one, what is kept and what is discarded?

level: middleimportance: should knowfreq 50%

basics

~20 s

A restart in place reuses the same container: same identity, same settings, same private writable layer and everything written into it. Only the process restarts, so memory, caches and in-flight work are lost. A replacement is a new container with a new empty writable layer.

open as a page

An ingester has been in a restart loop for six hours; a teammate proposes never restarting it — what does that change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It ends the churn and leaves the last failed instance in place to inspect, but it changes nothing about the cause and removes the automatic recovery that would have fixed a transient one. A restart loop is a symptom, not a diagnosis.

open as a page

How would you size the stop-to-kill grace period for a ledger service whose longest legitimate write takes about 90 seconds?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Size it from measurement: the drain pause that lets routing removal propagate, plus a high percentile of the longest legitimate unit of work, plus a margin. Ninety-second writes and a fifteen-second pause put the window near two minutes, not thirty seconds.

open as a page

Every instance of a service fails its readiness check at once because a shared downstream store is slow — what does the platform do?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It empties the routing set — every instance leaves at once, so even requests that never touch the slow store now fail, turning a partial degradation into a total outage. The processes keep running, unless the liveness check reads the same signal too.

open as a page

A renderer is force-killed at the end of every stop window — how do you tell a signal never delivered from one that was ignored?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Look for evidence that the application itself acted. The first line of its own shutdown path proves delivery; no such line, plus a first process that is a wrapper forwarding nothing or an application with no handler installed, points at a request that never arrived.

open as a page

A nightly reconciliation container always exits zero even on runs whose work failed, and its entry script runs the job then writes a summary — why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The container reports its first process's status, and that process is the entry script, not the job. The script's own status is its last statement's, so the summary write — which succeeds — is what the supervisor sees, and the job's failure is discarded.

open as a page

Your platform team wants one shutdown contract for every service on the estate - what do you mandate, and what do you leave to each team?

level: principalimportance: should knowfreq 34%

basics

~20 s

Mandate the ordering and the evidence: leave the routing set first, drain, refuse, finish, release, exit - plus a measured drain pause and a counter for work cut by the forced kill. Leave the window's length and what finishing means to each service.

open as a page

A worker holding claimed queue messages and an exclusive lease is asked to stop - what must it release before exiting?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Everything it holds on behalf of others: claimed messages returned so another consumer can take them, and the exclusive lease given up so a replacement can acquire it. A worker that just exits leaves both held until their timeouts expire.

open as a page

Should a platform standard put a minimal init in every image, or require each application to be its container's first process?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Neither is free. A mandated minimal init makes every image survive wrappers and stray children, at the cost of an extra process and an exit-status contract. Requiring applications to hold position one is cheaper but is a discipline that regresses in images nobody reviewed.

open as a page