skip to content

A cron-run Ruby log-rotation script logs its failures, but cron always sees status 0; what in the script decides the exit status?

level: seniorimportance: should knowfreq 35%

answer

  1. a handled error is not a failure
  2. falling off the end means 0
  3. raise SystemExit defaults to success
  4. only 8 bits reach the parent
  5. exit(ok) maps true/false to 0/1

basics

~20 s

Only how the program ends: finishing normally, even after a rescued and logged error, returns 0. The script must end with exit(false), exit 1 or abort, nothing may cancel that SystemExit or override it in at_exit, and statuses wrap past 255.

solid answer

~40 s

Cron and other supervisors only see the exit status, and Ruby derives it from how the program ends. A normal finish is 0, even if an error was rescued and logged along the way; an uncaught exception is 1; `exit(status)` and `abort` raise `SystemExit`, whose status becomes the process status unless something rescues it or an `at_exit` handler calls `exit` again. So the fix is structural: keep the work in a method that returns success, and end the script with one explicit `exit(ok)` (`true` is 0, `false` is 1) or `abort "reason"`. Watch two quieter traps: `raise SystemExit` with no status means success, and POSIX keeps only the low 8 bits, so `exit 256` reads as 0.

code

ruby · 12 lines
ruby
# rotate_logs.rb, run from cron
def rotate(dir)
  Dir.glob(File.join(dir, "*.log")).each do |path|
    File.rename(path, "#{path}.1")
  end
  true
rescue SystemCallError => e
  warn "rotate_logs: #{e.class}: #{e.message}"
  false
end

exit rotate(ARGV.fetch(0, "/var/log/app"))   # true -> 0, false -> 1

go deeper

for a junior

Recall that 0 means success and non-zero means failure, and that a rescued error does not change the status by itself.

for a middle

Explain that exit and abort raise SystemExit, that rescuing it cancels the exit, and how exit(true/false) and abort map to statuses.

for a senior

Diagnose a silently green job: check rescued errors without a final exit, cancelled SystemExit, at_exit overrides and status ranges, then reshape the script around one explicit exit and verify with echo $?.

for a principal

Set a status contract for every scheduled job: documented codes, one exit point, and alerting that keys on status rather than on log text.

## The contract with cron A scheduler cannot read your log; it reads the **exit status**. By convention 0 means success and anything else means failure, and alerting (a mail from cron, a failed job in a scheduler, a red check) keys off that number. A Ruby script therefore has to make sure that every failure path ends with a non-zero status. ## How Ruby decides the status | How the program ends | Status | |---|---| | main script finishes normally | 0 | | `exit` / `exit(true)` | 0 | | `exit(false)`, `abort(msg)` | 1 | | `exit(n)` | `n`, truncated to 8 bits by the OS | | uncaught exception other than `SystemExit` | 1 | | `raise SystemExit` with no status | 0 | | an `at_exit` handler calls `exit(m)` | `m`, replacing the earlier status | ## Trap 1: rescue, log, fall off the end ```ruby def rotate!(dir) Dir.glob(File.join(dir, "*.log")).each { |p| File.rename(p, "#{p}.1") } rescue SystemCallError => e warn "rotate_logs: #{e.message}" end rotate!("/var/log/app") # status 0 even when the rename failed ``` Once the exception is rescued it no longer exists as far as the exit status is concerned. Writing to `$stderr` with `warn` does not change the status, and neither does the error's class: an `Errno::EACCES` does not become status 13. The main script finishes normally and returns 0. ## Trap 2: something cancels the `SystemExit` `exit 1` and `abort` work by raising **`SystemExit`**. Any code between the call and the top level that rescues it, whether a `rescue SystemExit` or a very broad rescue clause, cancels the exit: execution continues after that `begin` block and, unless something else fails, ends with 0. Search the call path for rescues broad enough to catch it. ## Trap 3: an `at_exit` handler overrides the status A handler that calls `exit` or `exit 0`, perhaps written to guarantee a tidy shutdown message, replaces whatever status the program had, because its own `SystemExit` becomes the final one. Check handlers registered by the script and by the libraries it loads. ## Trap 4: statuses that look like failure but are not - `raise SystemExit` without arguments defaults to status `true`, which is 0; pass a status (`raise SystemExit.new(1, "rotation failed")`) or call `abort`. - On POSIX systems only the low 8 bits reach the parent: `exit 256` is seen as 0 and `exit 257` as 1. Keep your own failure codes small; shells reserve values from 126 up for their own meanings. ## The fixed shape 1. Keep the work in a method that returns `true`/`false` or a list of failures. 2. Rescue `SystemCallError` or other specific classes inside it, log with context, and record the failure. 3. End the script with exactly one statement that maps the result: `exit(ok)` or `abort "reason"`. 4. Verify by hand: run the script, check `$?` in the shell with `echo $?`, then force a failure (a read-only directory) and check again. ## Testing the contract Unit tests should call the method and assert on its return value rather than letting the script's `exit` run inside the test process. For an end-to-end check, run the script as a child process from the test and assert on the child's exit status. ## A review checklist for scheduled scripts - Is there exactly one place where the script decides its status, and does every failure path reach it? - Does any rescue between an `exit` call and the top level catch `SystemExit`? - Do any `at_exit` handlers call `exit`, and if so, do they preserve a failing status by checking `$!.success?` first? - Are custom status codes documented and kept below 126? - Does the job's alerting read the status rather than search the log text? A script that passes this checklist fails loudly: the scheduler sees a non-zero status, and the log explains why.

  • Why does `warn` in the rescue clause not make cron treat the run as failed?
    `warn` only writes to standard error. Cron may mail that output, but the exit status, which schedulers and monitors actually check, is decided by how the program ends, and a rescued error followed by a normal finish is status 0.
  • How would you find out which of these traps a silently green job has hit?
    Reproduce the failure locally and print `$?` after the run. If it is 0, check whether the error was rescued without a final `exit`, whether a rescue on the call path catches `SystemExit`, and whether an `at_exit` handler calls `exit`. A status of 0 after `exit` with a large number points to 8-bit truncation.

saying these in an interview costs you the question

  • Writing to $stderr is enough for cron to mark the run failed
  • An Errno exception's number becomes the process exit status
  • raise SystemExit with no arguments reports a failure
  • exit 1 cannot be cancelled once it has been called
  • Any integer passed to exit reaches the shell unchanged