A telemetry ingester is in a restart loop, but its logs come back nearly empty — where is the failing attempt's output?
answer
- fresh instance, fresh stream
- you are reading the wrong attempt
- the failure ended before you looked
- ask for the previous attempt's output
- usually one prior attempt is kept
basics
~20 sEach restart attempt is a new instance with a fresh output stream, so a request for the logs returns the attempt that just started, not the one that failed. Read the previous attempt's retained output instead.
solid answer
~40 sA restart does not continue the old instance; it creates a new one, and that new instance's standard output and error start empty. By the time you look, the current attempt may be seconds old and still starting, so its stream holds a start-up banner and nothing else — the error you want was printed by the attempt that already ended. Platforms keep the immediately previous instance's stream for exactly this reason and give you a way to ask for it, so the first move on any looping workload is to request the *previous* attempt's output rather than the current one. Retention normally covers only that one prior attempt, so a loop steadily overwrites its own evidence; anything you need to keep has to be copied off the host while it still exists.
go deeper
Recall that every restart attempt is a brand-new instance with an empty output stream, and that the failing attempt's output is retained separately and has to be asked for by name.
Explain why the current stream is nearly empty at any moment during a loop, and note that retention usually covers only the immediately previous attempt, so a loop overwrites its own evidence.
Show how you preserve the evidence before it is displaced: take the copy at first contact, read how the stream ends rather than only what it says, and distinguish the original failure from secondary ones later attempts introduce.
Frame it as an observability requirement rather than a debugging trick. A workload whose only failure evidence dies with the instance is undebuggable at fleet scale, whatever the underlying cause turns out to be.
## What a restart actually creates A restart loop is a sequence of separate attempts, not one instance having a long bad day. When the first process inside a container ends — on its own, or because it was ended from outside — that instance is over. Whatever supervises the workload consults its **restart policy**, and if the policy asks for another attempt, a **new instance** is created from the same image and the same spec. That new instance has its own process and, the part that matters here, its own **standard output and error**: the two streams that are the log contract for a container on every platform. Nothing is carried across that boundary. The new attempt's streams begin empty and fill from its first line of output onward. ## Why the current stream is nearly always empty Now look at the timing. The thing you want to read happened in the attempt that has **already ended**. The attempt you can address right now is the one that has **just begun**. If the ingester accepts data for forty seconds and then dies, and you ask for its logs at a random moment, you are on average twenty seconds into a fresh attempt that has printed a banner and a configuration line and has not reached its failure yet. Two conclusions are available, and people routinely draw the wrong one: - The right conclusion is **you are reading the wrong attempt**. - The wrong conclusion is **this workload does not log anything** — a belief people then act on by adding logging to a workload that was already telling them what was wrong, in a stream they never opened. ## Where the evidence lives | where | what it holds | how long it lasts | |---|---|---| | the current attempt's stream | the attempt that has not failed yet | until this attempt ends | | the retained previous-attempt stream | the failure you are actually looking for | until the next attempt replaces it | | a copy taken off the host | each attempt that was copied in time | as long as the destination keeps it | Platforms expose the middle row deliberately — the retained output of the **immediately previous** instance — because this failure mode is universal and nobody can debug a loop without it. The names and the mechanisms differ between platforms; the idea does not. The important limit is in the third column. Retention normally covers **one** prior attempt. A restart loop is therefore self-erasing: attempt 41 overwrites the record of attempt 40, and by the time a human is looking, the *first* failure is long gone. That first failure is often the most informative one, because later attempts can be failing for a secondary reason — a half-written file, a connection the previous attempt never released, a queue position already claimed — that has nothing to do with what started the loop. ## The order to read in 1. Ask for the **previous attempt's** retained output, not the current attempt's. 2. Read it to the end. The last lines written are the ones closest to the ending, and they are the ones that matter. 3. Notice *how* the stream ends. A message about shutting down and an abrupt stop mid-line mean very different things about how the attempt finished. 4. If the previous attempt's stream is genuinely uninformative, stop re-reading streams and change the experiment: arrange for the output to be captured somewhere durable before the next attempt overwrites it. ## What this does not tell you Being disciplined about *which* attempt you are reading is a precondition for a diagnosis, not a diagnosis. The retained stream tells you what the process printed. It does not tell you whether the process chose to end or was ended from outside, and it will be empty rather than misleading if the attempt never got as far as running the workload's own code at all. Those are separate subjects with their own answers. What this leaf owns is the discipline itself: **in a restart loop, the currently running instance is the least interesting one on the host.** One practical note on the shape of the problem. The delay between attempts usually grows as the loop continues, and that delay is also what gives you time to read. Early in a loop, attempts arrive seconds apart and the retained stream is replaced while you are still scrolling it. Later, when the delay has grown to minutes, the same request returns a stable answer. That is not a reason to wait for the loop to slow down — it is a reason to take a copy the first time you have one, because the comfortable reading window opens exactly when the evidence is at its oldest and least representative.
- Why does the current attempt of a looping ingester usually show only a start-up banner?Because that attempt is seconds old. Its streams began empty when the instance was created, and it has not yet run long enough to reach the point where the previous attempts failed. The banner is simply the first thing every attempt prints.
- You want the output of the attempt before the previous one. Why is it usually unavailable?Host-side retention normally keeps one prior attempt, so each new instance's stream displaces the record two steps back. Anything older survives only if a copy was taken off the host before it was displaced.
- The previous attempt's retained stream is empty too. What does that narrow down?It says the attempt produced no output before it ended — so either it never reached the workload's own code, or it was ended from outside before it printed anything. It does not by itself say which; both need evidence from outside the stream.
saying these in an interview costs you the question
- Assumes the running instance's stream contains a crash that already happened
- Concludes the workload logs nothing because the current attempt's output is empty
- Thinks a restart resumes the same instance and continues its output stream
- Expects every earlier attempt in the loop to still be retrievable from the host
- Tries to open an interactive session inside an instance that lives for seconds