skip to content

questions

4

When a nightly batch container stops, what does its exit code tell whatever supervises it?

level: juniorimportance: must knowfreq 72%

answer

  1. one number leaves the container
  2. zero or not zero
  3. the supervisor's only verdict at stop
  4. meaning beyond zero is the program's
  5. a stop must not look like a crash

basics

~20 s

The exit code is the container's own verdict on its run: zero means the work completed, any non-zero value means it did not. At stop time the supervisor branches on that one number, not on anything the job printed.

solid answer

~40 s

When the container's first process ends, the runtime records the number it returned and the container moves to a stopped state. By convention `0` means the work finished and any non-zero value means it did not; beyond that split, individual numbers mean only what the program documents, so nothing portable can be read into them. Whatever supervises the container acts on that one value: a run-to-completion job with `0` is recorded as done and not attempted again, while a non-zero status is recorded as a failed run, is what alerting fires on, and is what a restart policy is consulted about. The log stream is evidence for a human; nothing parses it to decide success. That is why a deliberate stop and a crash must not present the same number.

code

pseudocode · 9 lines
pseudocode
on container_stopped(exit_code):
    if exit_code == 0:
        record_run(succeeded)
        if workload_shape == run_to_completion:
            do_not_start_again()
    else:
        record_run(failed, exit_code)
        raise_alert(exit_code)
        consult_restart_policy()

go deeper

for a junior

Remember the split and who reads it: zero means the work completed, non-zero means it did not, and something outside the container acts on that number automatically.

for a middle

Explain that the container's status is its first process's status, that only the zero/non-zero split is portable, and that a job and a long-running service draw opposite conclusions from the same zero.

for a senior

Show the operational consequence: alerting and retry decisions are computed from this one value, so a run cut short mid-work that reports zero is a silent data gap nobody is paged for.

for a principal

Frame it as an interface other teams depend on: what your batch jobs promise by exiting zero, where partial progress is recorded instead, and how failures reach a human without anyone parsing log text.

## The one number a container leaves behind When the container's first process ends, the runtime records the number that process returned and moves the container into a stopped state. That number — the **exit code**, also called the **exit status** — is the whole of what the container says about its own run. Everything else it produced is evidence for a human: log lines, files it wrote, metrics it emitted. The exit code is the part the machinery acts on, automatically, with nobody reading anything. The convention has exactly two levels of meaning: - **Zero** means *the work I was given completed*. - **Any non-zero value** means *it did not*. - Past that split nothing is portable. A program may use different non-zero values to separate classes of failure — bad input, an unreachable dependency, nothing to do — but those meanings are that program's own convention and say nothing to a reader who does not know the program. So the honest reading is binary, plus a hint. Treat the zero/non-zero split as the contract, and any particular number as documentation the job's author owes you. ## What the supervisor does with it *Supervisor* here means whatever is watching the container and deciding what happens next: a single-host runtime applying a restart policy, a cluster scheduler, a batch system, a pipeline runner. Whichever it is, the stop event it receives carries the exit code, and the decision it makes is a branch on that value. | workload shape | exits zero | exits non-zero | |---|---|---| | nightly run-to-completion job | the run is recorded as complete, and the job is not started again for that run | the run is recorded as failed, alerting fires on it, and the restart policy decides whether another attempt follows | | long-running service | the process ended although nobody asked it to, which is still an unplanned stop; platforms differ on whether a zero exit is restarted | treated as a crash: counted as a failure and restarted under the policy | Two things follow from the table. First, the same number means different things to different workload shapes — for a job, zero is the goal; for a service that was supposed to keep serving, zero is merely a quieter surprise. Second, the supervisor never learns *how much* was done. One number cannot say "processed 4,000 of 10,000 records". If partial progress matters, the job has to record it somewhere durable itself. ## Why a deliberate stop and a crash must not look alike The contract is only useful if the two outcomes it distinguishes actually stay distinct. - A container asked to stop, which finishes what it was doing and ends cleanly, should report zero. That is the planned case, and it should not raise an alert. - A container whose work failed should report non-zero, so the run is counted as failed and somebody is told. - A container that was asked to stop **mid-work** is the trap. If it reports zero, the supervisor records a complete run that in fact processed half its input, and the gap is discovered days later by whoever notices the missing data. - The inverse is just as bad: a service that reports a failure every time it is asked to shut down turns routine operations into a stream of false incidents, and everyone learns to ignore the signal. - When the platform has to end a container forcibly because it would not stop, the recorded status is not one the program chose. How an operating system encodes that is a separate subject; what matters here is that it is distinguishable from a self-chosen zero. ## What the exit code is not - **Not a log.** No supervisor greps output for the word "error" to decide whether a run succeeded, and a job that reports failure only in its text is a job whose failures are invisible. - **Not a progress report.** It is complete or not complete, nothing in between. - **Not a judgement about a *running* container.** While the container is up, the questions are answered by checks, not by a status that does not exist yet. - **Not the last trace of the run.** The container's record, its status and its captured output stay readable until the container is removed. ## Getting it right in practice 1. Make the process that does the real work the one whose status becomes the container's status. 2. Return zero only for a run that genuinely finished the work, and non-zero for every run that did not. 3. Document any non-zero values the job distinguishes, next to the job, because nothing else gives them meaning. 4. Alert and report on the recorded status, never on log text — the status is the only part of the run that arrives in a form a machine can act on.

  • A job distinguishes bad input from an unreachable dependency by returning different non-zero values. How much can a supervisor do with that?
    Nothing on its own. To every supervisor the two are the same event: a failed run. The distinction is only usable where something was deliberately taught the job's own mapping — a wrapper, an alert rule or a runbook written next to the job. The portable part of the contract remains zero against non-zero.
  • A long-running service exits zero of its own accord in the middle of the night. Is that a success?
    No. For a service, finishing is itself the anomaly — nobody asked it to stop, so something ended it: a fatal condition handled too politely, a caught error that fell through to a clean exit, or a code path that returns instead of continuing to serve. Zero only means the process considered its own ending normal, not that the workload is healthy.
  • The nightly job processed 4,000 of 10,000 records before failing. What should the exit code say?
    Non-zero, because the work it was given did not complete. The code cannot carry the 4,000, so the job must record its progress somewhere durable — a checkpoint, a marker row, a written offset — and the next attempt reads that. Using zero to mean "partly done" makes an incomplete run indistinguishable from a good one.

saying these in an interview costs you the question

  • Says a non-zero exit code always means the platform killed the container
  • Thinks the supervisor reads the log text to decide whether the run succeeded
  • Believes every non-zero number carries a standard, platform-defined meaning
  • Exits zero after a failure so the container is not restarted, and calls that graceful
  • Cannot say what zero changes for a job as opposed to a long-running service
open as a page

A container is created, runs, stops and is later removed — what triggers each of those transitions?

level: middleimportance: must knowfreq 60%

basics

~20 s

Creation builds the container from its spec without running anything; a start request executes its first process, which is what running means; the container stops when that process ends, by itself or because it was asked to; removal is a separate request that destroys the record.

open as a page

When a stopped container is restarted in place rather than replaced with a fresh one, what is kept and what is discarded?

level: middleimportance: should knowfreq 50%

basics

~20 s

A restart in place reuses the same container: same identity, same settings, same private writable layer and everything written into it. Only the process restarts, so memory, caches and in-flight work are lost. A replacement is a new container with a new empty writable layer.

open as a page

A nightly reconciliation container always exits zero even on runs whose work failed, and its entry script runs the job then writes a summary — why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The container reports its first process's status, and that process is the entry script, not the job. The script's own status is its last statement's, so the summary write — which succeeds — is what the supervisor sees, and the job's failure is discarded.

open as a page