skip to content

A nightly reconciliation container always exits zero even on runs whose work failed, and its entry script runs the job then writes a summary — why?

level: seniorimportance: should knowfreq 45%

answer

  1. a success nobody earned
  2. which process's status is reported
  3. the wrapper is what exits
  4. the last statement decides
  5. save the status, then clean up

basics

~20 s

The container reports its first process's status, and that process is the entry script, not the job. The script's own status is its last statement's, so the summary write — which succeeds — is what the supervisor sees, and the job's failure is discarded.

solid answer

~50 s

The container's exit code is the exit code of its **first process**, and here that is the wrapper. A wrapper's status is whatever its last statement produced, or zero if it simply ran off the end; the status of a command in the middle is dropped as soon as the next statement runs. So the failing job returns a non-zero value, the summary write succeeds, and the container reports success. The supervisor then records a completed reconciliation, raises no alert, and schedules nothing again — a silent failure, which is worse than a loud one because the gap is only found when someone notices missing data. The fix is to make the work's status the container's status: run the job as the first process directly, or capture its status, do the cleanup, and exit with the saved value.

code

pseudocode · 4 lines
pseudocode
# entry script as written
status = run(reconcile_batch)   # returns 3 on a failed run
write_summary(status)           # writes one line, succeeds
# script ends here, so the container reports 0

go deeper

for a junior

Take away one rule: the container reports the status of its first process. If a script wraps the real work, the script's status is what leaves the container.

for a middle

Explain how the status gets dropped — a later statement, a caught failure, a pipeline — and how to keep it by capturing the work's value and exiting with it after cleanup.

for a senior

Name the failure mode and how you would find it: silent success, a job whose recorded failure rate is exactly zero, and the comparison between the job's own outcome and the status the container reported.

for a principal

Treat the status as an interface your organisation depends on: what every batch job is required to promise, who notices when a run is quietly wrong, and how long such a defect could survive unreported today.

## The first process is the one that speaks A container's exit code is not "the worst thing that happened inside it". It is one number from one process: the first process the runtime executed. Everything else the container ran is a child, and a child's status is reported to its parent, not to the outside world. If the first process is a wrapper script, the wrapper is the only thing the supervisor ever hears from. A script's own status is the status of the last statement it ran, or zero when it simply reaches the end with nothing to report. So this sequence: 1. run the reconciliation job — it fails and returns a non-zero value; 2. write a summary line — it succeeds; 3. reach the end of the script; ends with a container that reports success. Nothing malfunctioned. The wrapper did exactly what it says, and the contract with the supervisor was satisfied with the wrong number. ## The ways the status gets lost - **A later statement overwrites it.** Cleanup, a summary, a notification, a final log line — anything after the real work becomes the status the container reports. - **An always-succeed clause.** Someone appends a clause that turns any failure into success, usually to stop a container restarting, and every future failure is invisible. - **Catching the failure and logging it.** The wrapper notices the non-zero value, prints a message, and carries on. The message is text; text is not the contract. - **A pipeline.** When the work's output is piped into another command, the status that survives is commonly the *last* stage's, not the failing stage's — so piping a failing job into a formatter reports the formatter's success. - **A background start.** The wrapper starts the work in the background and returns, and the container's run ends before the work does. ## What it actually costs | what the supervisor believes | what happened | |---|---| | the nightly run completed | the job failed partway and wrote nothing | | no alert is warranted | nobody is told, on any night | | no further attempt is needed | the missed run is never retried | | the schedule is healthy | the failure rate reads as zero while the data drifts | This is the specific failure mode worth naming in an interview: **silent success**. A crash loop is loud and gets attention within minutes. A job that fails while reporting success can run wrong for months, and the discovery is usually a business question — a report that does not reconcile — rather than an alert. ## The fix 1. **Prefer no wrapper.** If the real work can be the container's first process, its status is the container's status by construction and none of this can happen. 2. **If a wrapper is needed, save the status.** Capture the work's status immediately, run whatever must run afterwards, and exit the wrapper with the saved value. Cleanup then still happens, and it no longer decides the verdict. 3. **Make an interrupted run distinguishable.** A run cut short mid-work is not a completed run. Either end with a non-zero status, or have the job record its own completion durably so that the supervisor's verdict and the job's record can be checked against each other. 4. **Never let text be the alerting path.** If the only evidence of failure is a line in the output, failures are only found by whoever happens to read it. ## Proving it is the wrapper and not the job - Read the recorded status of the failed runs. If it is zero on every single run, including ones you know failed, the status is being manufactured rather than reported. - Compare the job's own logged outcome against the container's recorded status for the same run. A disagreement localises the fault to what sits between them. - Run the work as the first process, without the wrapper, and see whether a failure now surfaces as a non-zero status. - Have the wrapper log the value it captured, so the number it *saw* and the number it *reported* are both visible. The general principle sits above this one scenario: whatever sits between the real work and the supervisor is part of the contract, and anything that rewrites the status on the way out has quietly redefined what "a successful run" means for everyone downstream of it.

  • The same job is asked to stop mid-run during a planned host change and its wrapper exits zero. Why is that just as bad?
    Because a run that processed half its input is recorded as a completed reconciliation. The supervisor has no other input, so nothing is retried and nothing is flagged, and the partial result is treated as the night's answer. An interrupted run should either report non-zero or write its own durable completion marker that something later checks.
  • Someone argues that exiting zero on failure is deliberate, because a non-zero status would make the container restart in a loop. Is that reasonable?
    No — it treats the reporting contract as a restart switch. Whether a failed job is attempted again is a property of its restart policy, and that is where to change it. Falsifying the status to control restarts also destroys alerting, failure counts and every report built on them.
  • How would you catch this class of defect across many jobs rather than one?
    Look for jobs whose recorded status has been zero on every run since they were created — a failure rate of exactly zero over a long period is a sign of a job that cannot report failure, not a perfect job. Then check each one for a wrapper that runs something after the real work.

saying these in an interview costs you the question

  • Says the container reports the worst status of everything it ran inside
  • Appends an always-succeed clause so the container stops restarting
  • Insists the job's failure message in the output is enough to alert on
  • Thinks the longest-running command in a wrapper decides the status
  • Treats a run cut short mid-way as completed because it exited zero