skip to content

In a Ruby service, which non-StandardError exceptions should a top-level crash reporter catch, and which must it let end the process?

level: seniorimportance: should knowfreq 30%

answer

  1. report, then re-raise
  2. shutdown is not a crash
  3. SystemStackError is rescuable
  4. NoMemoryError is preallocated
  5. OOM kill is SIGKILL

basics

~10 s

Let SignalException, Interrupt and SystemExit pass unreported, since they are shutdown. Report ScriptError, SystemStackError and NoMemoryError as crashes, allocating little, then re-raise so the process still terminates.

solid answer

~40 s

At the outermost edge, a handler may rescue `Exception` so that crashes outside `StandardError` get reported, but it must distinguish control flow from failure. `SignalException` (with `Interrupt`) and `SystemExit` are how Ruby shuts down; re-raise them untouched, or a clean `exit 0` and every deploy's TERM show up as crashes. `ScriptError`s (a `LoadError` from a lazily loaded file, a `SyntaxError` from `eval` or `load`) and `SystemStackError` from runaway recursion are genuine crashes to report. `NoMemoryError` is raised from a preallocated object when allocation fails, so the reporter should do as little allocating as possible. In every case the handler re-raises. Note that a container's out-of-memory kill is SIGKILL: Ruby sees no exception at all.

code

ruby · 10 lines
ruby
def run_process
  yield
rescue SignalException, SystemExit
  raise                                        # shutdown paths, not crashes
rescue Exception => e # rubocop:disable Lint/RescueException
  $stderr.write("crash: ", e.class.name, "\n")  # keep allocation low
  raise
end

run_process { main_loop }

go deeper

for a junior

Know that a top-level handler is the one place rescue Exception may appear, and that it must re-raise.

for a middle

List the non-StandardError classes and sort them into shutdown (SignalException, SystemExit) and crashes (ScriptError, SystemStackError, NoMemoryError).

for a senior

Show the operational detail: clause order, low-allocation reporting for NoMemoryError, trimmed stack-overflow traces, and SIGKILL out-of-memory kills that Ruby never sees.

for a principal

Own the policy: one audited catch-all per process, Lint/RescueException disabled only there, and crash alerts that exclude normal shutdowns.

## Where a catch-all belongs Most Ruby code should rescue **`StandardError`** or narrower classes. There is one legitimate exception to that rule: the **outermost frame of a process**, such as the method that runs a worker's main loop or a script's entry point. There, a handler may rescue **`Exception`** so that failures outside `StandardError` still reach logs or an error tracker before the process dies. Two rules make that safe: 1. **Classify** what was caught: shutdown or crash. 2. **Re-raise** in every case, so Ruby's own termination behaviour, including `ensure` blocks further out and the exit status, is preserved. ## The non-StandardError classes, one by one | Class | Meaning | Handler action | |---|---|---| | `SignalException`, `Interrupt` | TERM, HUP, Ctrl-C: someone asked the process to stop | re-raise without reporting | | `SystemExit` | `exit` was called | re-raise without reporting | | `ScriptError` (`LoadError`, `SyntaxError`, `NotImplementedError`) | the code is broken or incomplete | report, re-raise | | `SystemStackError` | stack overflow, "stack level too deep" | report, re-raise | | `NoMemoryError` | memory allocation failed | report minimally, re-raise | ### Shutdown is not a crash `SignalException` and its subclass `Interrupt` are raised in the main thread when the process receives signals such as TERM or INT. `SystemExit` is raised by `exit`. Reporting them turns every deploy, every Ctrl-C and every successful `exit 0` into an alert, and alert fatigue then hides the real crashes. List them in their own rescue clause before the catch-all and re-raise. ### Broken code `LoadError` and `SyntaxError` usually surface at boot, but code that loads files lazily or evaluates strings can raise them hours later. They are deploy defects: report them with the file name so the release can be rolled back. ### Stack overflow `SystemStackError` inherits directly from `Exception`, so a bare `rescue` in application code never saw it. By the time the top-level handler runs, the stack has unwound, so reporting is safe. The backtrace is very long and repetitive; trimming it before sending keeps the report readable. ### Memory exhaustion When CRuby fails to allocate memory it raises **`NoMemoryError`** using an exception object created at startup (message `failed to allocate memory`), precisely because allocating a new one might fail. A reporter that builds large strings or JSON payloads at this point may fail again. Write a short fixed line to standard error and re-raise. ## What a Ruby handler can never see - **SIGKILL.** No process can catch it. Container runtimes enforce memory limits by having the kernel kill the process with SIGKILL, so an out-of-memory kill in a container produces no `NoMemoryError` and no Ruby exception at all. Diagnose those from the orchestrator's exit reason and memory metrics. - **The internal `fatal` class**, which Ruby uses when it must exit immediately; application code does not rescue it. ## Lint and review RuboCop's `Lint/RescueException` flags any clause that names `Exception`, even one that re-raises. For the single sanctioned top-level handler, disable the cop inline with a comment saying why, so every other `rescue Exception` in the codebase is still reported. In review, check the three properties that make the handler safe: - shutdown classes rescued first and re-raised unreported; - a reporter that allocates little and cannot itself raise past the re-raise; - a final `raise`, so the process exits with the original exception. ## A worked flow Consider a worker process started by a supervisor, with the handler above around its main loop: 1. A deploy sends TERM. Ruby raises `SignalException` in the main thread, the first clause re-raises it, `ensure` blocks run and the process exits without an alert. 2. A job triggers infinite recursion. `SystemStackError` reaches the handler, one short crash line is written, the exception is re-raised and the supervisor restarts the worker. 3. A lazily loaded file was left out of the release. `LoadError` is reported with its path and re-raised, so the crash is visible in the release's error rate. 4. The container exceeds its memory limit. The kernel sends SIGKILL; nothing in Ruby runs, and the only evidence is the orchestrator's record of the kill.

  • Why is reporting NoMemoryError unreliable?
    It is raised when the interpreter cannot allocate memory, from an exception object Ruby preallocated for that purpose. A reporter that formats payloads or opens connections needs memory too and may fail. Keep the handler to a short fixed write, re-raise, and rely on external monitoring for memory trends.
  • Why must the shutdown clause come before rescue Exception?
    Ruby tries rescue clauses in order and uses the first whose class matches. `Exception` matches everything, so if it came first, `Interrupt` and `SystemExit` would be reported as crashes and the dedicated clause would never run.
  • Does rescuing SystemStackError make deep recursion safe?
    No. It lets you report the failure, but the operation that overflowed did not finish. The fix is to bound the recursion or make it iterative; continuing after a stack overflow just hides a defect that will recur with the same input.

saying these in an interview costs you the question

  • Reporting SystemExit and Interrupt as crashes is harmless
  • A container hitting its memory limit raises NoMemoryError in Ruby
  • rescue Exception is fine as long as the error is logged and the loop continues
  • SystemStackError is a StandardError, so a bare rescue already reports it
  • Re-raising inside the clause stops Lint/RescueException from flagging it