skip to content

You're designing async exception handling for a platform with many fire-and-forget @Async tasks (emails, audit writes, webhooks). How do you make failures observable and recoverable without falling into the 'silently swallowed' trap?

level: principalimportance: should knowfreq 35%

answer

  1. custom handler: log + metric + alert/dead-letter
  2. return-type policy: void vs must-observe CompletableFuture
  3. unobserved future = silent loss (worst trap)
  4. TaskDecorator: MDC/trace/SecurityContext propagation
  5. bounded queue + rejection policy + retry/replay

basics

~20 s

Don't rely on the default log-only handler. Register a custom AsyncUncaughtExceptionHandler (via AsyncConfigurer) that logs with context, emits metrics, and alerts or dead-letters. For work needing results or retries, return CompletableFuture and handle failures explicitly, and propagate trace/MDC context onto async threads.

solid answer

~50 s

I treat async failure as a first-class design concern. First, replace SimpleAsyncUncaughtExceptionHandler with a custom AsyncUncaughtExceptionHandler via AsyncConfigurer — it logs with method+args context, increments a failure metric (per method), and for critical tasks emits an alert or writes a dead-letter record for replay. Second, I pick return types deliberately: truly fire-and-forget stays void (covered by the handler); anything whose outcome matters returns CompletableFuture with explicit exceptionally/handle so failures aren't lost in an unobserved future. Third, I propagate context — a TaskDecorator copies MDC/trace IDs and SecurityContext onto the worker thread so handler logs are correlatable. Fourth, I configure the executor deliberately (bounded queue, named threads, sensible rejection policy) and add retries via Spring Retry or idempotent re-enqueue for transient failures. Finally, monitoring: dashboards/alerts on the async-error metric, not log-grepping. The anti-patterns I guard against: unobserved CompletableFutures, catch-and-drop handlers that kill even the default log, and assuming caller transactions roll back.

code

java · 31 lines
java
@Configuration
@EnableAsync
public class AsyncConfig implements AsyncConfigurer {

    private final MeterRegistry meters;
    private final DeadLetterStore deadLetters;
    // ctor injection omitted

    @Override
    public Executor getAsyncExecutor() {
        ThreadPoolTaskExecutor exec = new ThreadPoolTaskExecutor();
        exec.setCorePoolSize(8);
        exec.setMaxPoolSize(16);
        exec.setQueueCapacity(500);                 // bounded: surface overload
        exec.setThreadNamePrefix("async-");
        exec.setRejectedExecutionHandler(new ThreadPoolExecutor.CallerRunsPolicy());
        exec.setTaskDecorator(new ContextCopyingDecorator()); // MDC + SecurityContext
        exec.initialize();
        return exec;
    }

    @Override
    public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() {
        return (ex, method, params) -> {
            String m = method.getDeclaringClass().getSimpleName() + "." + method.getName();
            LoggerFactory.getLogger(m).error("async-failure args={}", Arrays.toString(params), ex);
            meters.counter("async.errors", "method", m).increment();
            deadLetters.record(m, params, ex);      // enable replay
        };
    }
}

go deeper

for a junior

Knows the default logs and that's usually not enough.

for a middle

Can implement a custom handler with metrics and pick return types.

for a senior

Adds context propagation, bounded executors, and retry/dead-letter thinking.

for a principal

Defines the end-to-end async-failure policy (observability + recoverability + conventions + monitoring) and enforces it across teams.

## Framing The default behavior (void → `SimpleAsyncUncaughtExceptionHandler` logs ERROR; futures capture their own error) is fine for a demo but leaves two production gaps: **observability** (log-grepping doesn't scale, and unobserved futures log nothing) and **recoverability** (nothing retries or replays a failed side effect). A principal-level answer builds an explicit policy. ## 1. A real `AsyncUncaughtExceptionHandler` (for void tasks) Register it via `AsyncConfigurer.getAsyncUncaughtExceptionHandler()`. It should: - Log at ERROR with **structured context**: method name, args (scrubbed of secrets/PII), correlation/trace id. - **Emit a metric** (e.g. Micrometer `Counter` tagged by method) so alerting is on metrics, not logs. - For critical tasks, **alert** and/or **dead-letter**: persist a record (task type, payload, error, timestamp) to a table/queue for later inspection and replay. ```java @Override public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() { return (ex, method, params) -> { String m = method.getDeclaringClass().getSimpleName() + "." + method.getName(); log.error("async-failure method={} args={}", m, scrub(params), ex); meterRegistry.counter("async.errors", "method", m).increment(); deadLetterStore.record(m, params, ex); // for replay }; } ``` ## 2. Deliberate return-type policy - **void** for pure side effects the caller is indifferent to → handler covers failures. - **`CompletableFuture<T>`** when the result or failure must be consumed, aggregated, or fed back into the request. Always attach `exceptionally`/`handle`/`whenComplete` — an **unobserved `CompletableFuture` failure is swallowed with no handler and no log**, the single worst trap. Establish a team convention: "if you return a future, someone must consume it." ## 3. Context propagation Handlers and async bodies run on **worker threads**, so MDC, trace IDs, and `SecurityContext` are absent by default — making failure logs uncorrelatable. Fix with a `TaskDecorator` on the executor that snapshots and restores `MDC` and `SecurityContext` around the task. (Spring 6 / Micrometer `ContextPropagation` and `ContextSnapshot` formalize this.) ## 4. Executor and back-pressure Failure handling includes **rejection**: a `ThreadPoolTaskExecutor` with an unbounded queue hides overload; a bounded queue + explicit `RejectedExecutionHandler` (e.g. `CallerRunsPolicy` or a custom one that dead-letters) surfaces it. Name threads for diagnosability. ## 5. Retries / recovery Transient failures (webhook 503, SMTP blip) want retry, not just logging: `@Retryable` (Spring Retry) inside the async method, or re-enqueue from the dead-letter store with backoff. Ensure the task is **idempotent** so replays are safe. ## 6. Monitoring Alert on the `async.errors` metric and dead-letter depth — not on grep. Track per-method error rates. ## Anti-patterns to call out - Unobserved `CompletableFuture` → silent loss. - A custom handler that catches and returns without logging → destroys the only default signal. - Assuming the caller's `@Transactional` rolls back on async failure (it doesn't). - Relying on the handler for future-returning methods (it's void-only). - Blocking on `get()` right after the call, making async pointless. ## When to use what Fire-and-forget with a robust handler for non-critical side effects; futures + explicit recovery for result-bearing or business-critical async; retry/dead-letter for anything that interacts with flaky external systems.

  • What's the single biggest silent-failure trap and how do you prevent it organizationally?
    An @Async method returning a CompletableFuture whose failure nobody observes — no handler fires and nothing is logged. Prevent it with a team convention (every returned future must be consumed), lint/review checks, and preferring void+handler when the caller genuinely won't consume the result.
  • How do you make async-failure logs correlatable with the originating request?
    Add a TaskDecorator to the executor that snapshots MDC, trace/correlation IDs, and SecurityContext on the caller thread and restores them on the worker thread around task execution (Spring 6 / Micrometer ContextSnapshot formalizes this).
  • Where would you add retry for a flaky webhook @Async task?
    Wrap the external call with Spring Retry (@Retryable + backoff) inside the async method for transient failures, and fall back to a dead-letter store for replay on exhaustion. Ensure idempotency so retries/replays are safe.

saying these in an interview costs you the question

  • Relying solely on the default log-only handler in production
  • Leaving CompletableFutures unobserved
  • Writing a handler that catches and drops without logging or metrics
  • Ignoring context propagation so failure logs have no trace id
  • Unbounded executor queue that hides overload

context