A long-running Node service hits an unhandled promise rejection in production. Should the process crash, as Node does by default, or should you install a handler that logs and keeps serving? How would you decide?
answer
- you cannot describe the state it left
- who restarts you, and how fast
- the hook changes policy, not just logging
- one bug versus slow corruption
- rare enough to be worth an alert
basics
~20 sDefault to crashing: an unhandled rejection means an error path nobody designed for, so the process state is unknown. Register a hook only to add context and shut down cleanly with a non-zero exit code, letting a supervisor restart a fresh process.
solid answer
~60 sAn unhandled rejection is by definition a failure your code did not anticipate, so you cannot reason about what state it left behind — a half-written record, a lease never released, a request that will never get a response. Continuing to serve from that state is how one bug becomes a slow corruption. Node's default `throw` policy is therefore the right baseline for a service that runs under a supervisor and can be replaced in seconds. Where a hook earns its place is observability, not survival: register `process.on('unhandledRejection', ...)` to attach request context and flush the error to your reporter, then stop accepting new work, let in-flight requests drain briefly, and exit non-zero. The failure mode to avoid is a hook that only logs, because registering any listener disables the default crash and silently converts fail-fast into fail-quiet. The complementary discipline is making the hook rare: every intentional fire-and-forget call gets its own `.catch()` at the call site, so anything reaching the global hook is a genuine defect worth an alert.
code
javascript · 9 linesprocess.on('unhandledRejection', (reason) => {
// Observability: add context and flush to the reporter.
logger.fatal({ reason }, 'unhandled rejection - shutting down');
// Policy: stay fail-fast, but drain first.
process.exitCode = 1;
server.close(() => process.exit(1));
setTimeout(() => process.exit(1), 5000).unref(); // hard deadline
});go deeper
Know that Node exits on an unhandled rejection by default and that this is intentional, not a bug — the error reached a path nobody wrote a handler for.
Explain why unknown process state, rather than the error itself, is what makes continuing risky, and that registering an unhandledRejection listener replaces Node's default exit behaviour.
Design the hook: attach request context, flush to the reporter, stop accepting new work, drain briefly with a hard deadline, and exit non-zero so the supervisor replaces the instance.
Own the tradeoff across a fleet — the preconditions that make crash-only safe, the stateful cases where it is not, restart-storm risk, and the call-site discipline that keeps the global hook a rare, alert-worthy defect signal.
## Start from what the signal means An unhandled rejection is not "an error occurred". Errors occur constantly in a healthy service and are handled at their call sites. An unhandled rejection means an error occurred **on a path where nobody wrote a handler** — which is a much stronger statement. It tells you the code reached a state its author never considered, and it tells you nothing about what was left half-done. That is why the question is not really "crash or log". It is: can you characterise the state the process is in? If you cannot — and by construction you cannot, because you did not anticipate this path — then continued execution is unbounded risk. ## The case for crashing Node's default since v15 is `throw`: raise the rejection as an uncaught exception, print it, exit non-zero. For a stateless request-handling process supervised by an orchestrator, this is close to ideal: - The blast radius is bounded to the requests in flight on that instance. - Restart cost is seconds, and a fresh process has provably clean state. - The failure is loud, which is what gets it fixed. The preconditions are worth naming explicitly, because they are what make crashing safe: more than one instance behind a load balancer, a supervisor that restarts automatically, restart-loop protection with backoff, and no important state that lives only in this process's memory. ## The case against reflexive crashing Crashing is a poor answer when those preconditions fail. A single-instance job runner mid-way through a batch, a process holding an in-memory queue that has not been persisted, a desktop or CLI tool where exiting destroys the user's work — in those cases exiting *is* the data loss. The honest response there is not to install a swallow-all hook; it is to fix the architecture so the state is not exclusively in memory, and until then to make the hook's shutdown path do the persisting. There is also the restart-storm concern: if the rejection is triggered by a request pattern that every instance receives, crash-on-rejection turns a bug into a fleet-wide outage. That is a real risk, and it is an argument for rate-limiting restarts and for canary deploys — not an argument for continuing to serve from unknown state. ## The trap nobody expects ```js process.on('unhandledRejection', (reason) => { logger.error({ reason }, 'unhandled rejection'); }); ``` This looks like pure observability. It is not: Node's `throw` policy applies **only when no `unhandledRejection` listener is registered**. Adding this hook silently changes the service's crash policy. Six months later nobody remembers that the three-line logging improvement is why the service now runs indefinitely in corrupted states. The fix is to make the policy explicit inside the hook: ```js process.on('unhandledRejection', (reason) => { logger.fatal({ reason }, 'unhandled rejection — shutting down'); process.exitCode = 1; server.close(() => process.exit(1)); setTimeout(() => process.exit(1), 5000).unref(); // hard deadline }); ``` Now the hook adds context and a graceful drain without giving up fail-fast. Node also lets you state the policy at the process level with `--unhandled-rejections=<mode>`, whose modes include `throw`, `strict`, `warn`, `warn-with-error-code` and `none`; choosing anything other than the default should be a documented decision, not a copied flag. ## Keep the hook a detector, not a net The global hook is only useful if reaching it is rare and always means a bug. That requires the complementary discipline at every call site: - Anything you expect to fail is handled where it fails, with the context to retry or degrade. - Every deliberate fire-and-forget call carries its own `.catch()` that logs and moves on. - Lint rules flag promises that are neither awaited, returned, nor caught. If those are in place, an entry in the global hook is a page-worthy event. If they are not, the hook fills with routine noise, everyone stops looking at it, and you have neither observability nor safety. ## How to present the decision A strong answer states the default (crash), names the preconditions that make it safe (multiple instances, supervisor, no memory-only state), names the exceptions (single-instance stateful workers, user-facing tools), and — most tellingly — flags the listener-disables-the-crash trap, because that is the part teams discover the hard way. Framing it as "observability belongs in the hook, survival does not" keeps the two concerns from being conflated.
- What has to be true about the deployment before crash-on-rejection is a safe default?More than one instance behind a load balancer, a supervisor that restarts automatically, backoff or rate limiting so a repeated trigger cannot become a restart storm, and no important state that exists only in that process's memory. Where those do not hold — a single-instance batch worker, a CLI tool — exiting is itself the data loss, and the architecture is what needs fixing.
- Why is a logging-only unhandledRejection hook considered a policy change rather than pure observability?Because Node applies its default `throw` behaviour only when no `unhandledRejection` listener is registered. Adding any listener takes over the policy, so a hook that merely logs converts crash-on-rejection into run-forever-in-unknown-state. If you register one, it must decide explicitly whether the process ends, and set a non-zero exit code when it does.
- How do you keep the global hook meaningful over time?By making it rare. Handle expected failures at their call sites, give every deliberate fire-and-forget call its own `.catch()`, and lint for promises that are never awaited, returned, or caught. Then every entry in the global hook is a genuine unanticipated path and deserves an alert; without that discipline it fills with routine noise and everyone stops reading it.
saying these in an interview costs you the question
- Says always keep serving because uptime matters most
- Adds a logging hook without noticing it disables the crash
- Treats the global hook as the service's error-handling layer
- Claims the process state is fine because the error was logged
- Ignores whether a supervisor exists before choosing to crash