skip to content

How do you decide which exceptions to catch locally versus let propagate, and where should broad 'catch-all' handlers live in a layered application?

level: seniorimportance: should knowfreq 45%

answer

  1. Catch only where you can act
  2. Recover / translate / add context — else propagate
  3. Don't log-and-rethrow at every layer (dup logs)
  4. One broad catch at the boundary: log once + safe response
  5. Never leak stack traces to external clients

basics

~20 s

Catch an exception only at the level that can actually do something useful about it. Let everything else bubble up. Put one broad catch-all at the application's outer edge — like a request handler — to log any failure and return a safe response.

solid answer

~50 s

The guiding principle is 'catch where you can act'. A catch only earns its place if this layer can recover (retry, fall back, return a default) or must translate the exception to its own abstraction. Otherwise, let it propagate — adding a local catch that just logs-and-rethrows at every layer produces duplicate logs and noise. Concentrate handling at two kinds of sites: (1) specific recoverable points deep in the code where a sensible fallback exists, and (2) a single broad boundary at the outer edge — the HTTP/RPC handler, message-consumer loop, or thread root — whose job is to catch any otherwise-unhandled failure, log it once with full context and a correlation id, and convert it to a safe error response so one request can't crash the process. Inner layers throw meaningful (often translated) exceptions; the boundary maps them to status codes/error payloads. This keeps business logic clean and makes failures observable and consistent.

go deeper

for a junior

Understands that not every method needs a try/catch and that errors can be allowed to bubble up.

for a middle

Places specific recovery catches where a fallback exists and knows there should be a top-level handler, without over-catching in between.

for a senior

Articulates 'catch where you can act', concentrates broad handling at boundaries, avoids duplicate logging, and prevents leaking internals to clients.

for a principal

Designs the cross-cutting error-handling architecture (boundary handlers, correlation ids, fail-fast vs. degrade policy, exception-to-status mapping) consistently across services and teams.

## The core decision rule For any potential catch site, ask: **can this layer do something meaningful about the failure?** Meaningful means one of: - **Recover** — retry, use a cached/default value, switch to a fallback, skip an optional step. - **Translate** — convert the exception into one that fits this layer's abstraction (e.g. `SQLException` → `RepositoryException`) so callers don't leak lower-level details. - **Add essential context and rethrow** — e.g. attach the entity id, then rethrow. If none of these apply, **do not catch here** — let the exception propagate. A catch that merely logs and rethrows the same thing at every layer creates **duplicate log entries** for one failure and clutters the code without adding value. ## Why local handling should be specific and rare Deep in the code, catches should be **narrow** (specific types) and tied to a concrete recovery. Example: a config loader catches `FileNotFoundException` and returns defaults — that is a real recovery at exactly the right place. It does *not* catch `IOException` broadly, because a disk read error mid-file is not something it can sensibly paper over. ## The boundary catch-all Every long-running program has an **outer boundary** where execution enters from the outside world: - An HTTP/RPC **request handler** (one catch per request). - A **message-consumer** loop (one catch per message). - A **scheduled-job** runner (one catch per run). - A raw **thread's run()** or an executor task. At these boundaries a **broad** `catch (Exception e)` is correct and necessary: its job is to ensure that *any* unhandled failure does not kill the whole server or silently drop work. It should: 1. **Log once**, with the full stack trace and context (request id / correlation id, user, operation). 2. **Convert** the failure into a safe outcome — an HTTP 500/4xx with a sanitized message, a dead-letter for the message, a failed-job record. 3. **Not leak internals** — never return a raw stack trace or DB error to an external client (information disclosure). In frameworks you usually express this declaratively (e.g. a Spring `@ControllerAdvice` / `@ExceptionHandler`, a servlet filter, a `Thread.UncaughtExceptionHandler`) rather than hand-writing try/catch in every endpoint. ## Putting the layers together ``` Client │ (HTTP) ▼ [Boundary] @ControllerAdvice: catch (Exception) -> log once + map to 500/4xx ▼ [Service] may catch + recover (retry, fallback) OR translate; else propagate ▼ [Repository] catch SQLException -> throw RepositoryException(msg, cause); else propagate ▼ [Driver] throws SQLException ``` Inner layers throw **meaningful** exceptions (translated where it helps); the boundary is the one place that decides the user-facing outcome. This gives you: clean business logic (no defensive catch noise), single-point logging (no duplicates), consistent error responses, and full diagnosability via preserved causes. ## Common pitfalls - **Catch-and-log at every layer** → duplicate logs, hard to tell how many failures actually occurred. - **No boundary handler** → an unexpected `RuntimeException` crashes the thread or returns an ugly raw trace to the client. - **Boundary catches `Throwable`** → it may swallow `Error`s (OOM) it can't handle; prefer `Exception`, and let `Error` crash or be handled very deliberately. - **Leaking internal messages** to clients → security/information-disclosure issue; log details server-side, return a generic message + a correlation id the user can quote. - **Recovering by returning null** instead of a typed result → pushes a hidden failure to the caller. ## Fail-fast vs. degrade-gracefully The boundary policy also encodes a philosophy: for *programmer errors* (NPE, illegal argument) you generally want **fail-fast** — surface them so they get fixed. For *expected operational failures* (a downstream timeout) you may want **graceful degradation** — fallback or partial response. Senior+ engineers choose per boundary which stance applies.

  • How do you avoid duplicate log entries for a single failure across layers?
    Adopt a 'log once, at the boundary' rule: inner layers translate/rethrow without logging (or log at debug only), and the outermost handler logs the full chain once with a correlation id. Logging the same exception at each catch is the usual cause of duplicates.
  • Why should a boundary handler avoid returning the raw stack trace to a client?
    It is an information-disclosure risk: stack traces reveal class names, library versions, file paths, and SQL, helping an attacker. Return a generic message plus a correlation id, and keep full details in server-side logs.
  • Where would you put a catch-all in a worker that processes a message queue?
    Around the per-message processing call: catch broadly, log once with the message id, then either dead-letter or nack the message so one poison message doesn't stop the consumer loop or get silently lost.

saying these in an interview costs you the question

  • try/catch logging at every single layer
  • No top-level handler, so RuntimeExceptions crash threads or leak traces
  • Boundary handler catching Throwable instead of Exception
  • Returning raw exception messages/stack traces to API clients
  • Using catch to return null and hide the failure from callers

context