You're designing async exception handling for a platform with many fire-and-forget @Async tasks (emails, audit writes, webhooks). How do you make failures observable and recoverable without falling into the 'silently swallowed' trap?
answer
- custom handler: log + metric + alert/dead-letter
- return-type policy: void vs must-observe CompletableFuture
- unobserved future = silent loss (worst trap)
- TaskDecorator: MDC/trace/SecurityContext propagation
- bounded queue + rejection policy + retry/replay
basics
~20 sDon't rely on the default log-only handler. Register a custom AsyncUncaughtExceptionHandler (via AsyncConfigurer) that logs with context, emits metrics, and alerts or dead-letters. For work needing results or retries, return CompletableFuture and handle failures explicitly, and propagate trace/MDC context onto async threads.
solid answer
~50 sI treat async failure as a first-class design concern. First, replace SimpleAsyncUncaughtExceptionHandler with a custom AsyncUncaughtExceptionHandler via AsyncConfigurer — it logs with method+args context, increments a failure metric (per method), and for critical tasks emits an alert or writes a dead-letter record for replay. Second, I pick return types deliberately: truly fire-and-forget stays void (covered by the handler); anything whose outcome matters returns CompletableFuture with explicit exceptionally/handle so failures aren't lost in an unobserved future. Third, I propagate context — a TaskDecorator copies MDC/trace IDs and SecurityContext onto the worker thread so handler logs are correlatable. Fourth, I configure the executor deliberately (bounded queue, named threads, sensible rejection policy) and add retries via Spring Retry or idempotent re-enqueue for transient failures. Finally, monitoring: dashboards/alerts on the async-error metric, not log-grepping. The anti-patterns I guard against: unobserved CompletableFutures, catch-and-drop handlers that kill even the default log, and assuming caller transactions roll back.
code
java · 31 lines@Configuration
@EnableAsync
public class AsyncConfig implements AsyncConfigurer {
private final MeterRegistry meters;
private final DeadLetterStore deadLetters;
// ctor injection omitted
@Override
public Executor getAsyncExecutor() {
ThreadPoolTaskExecutor exec = new ThreadPoolTaskExecutor();
exec.setCorePoolSize(8);
exec.setMaxPoolSize(16);
exec.setQueueCapacity(500); // bounded: surface overload
exec.setThreadNamePrefix("async-");
exec.setRejectedExecutionHandler(new ThreadPoolExecutor.CallerRunsPolicy());
exec.setTaskDecorator(new ContextCopyingDecorator()); // MDC + SecurityContext
exec.initialize();
return exec;
}
@Override
public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() {
return (ex, method, params) -> {
String m = method.getDeclaringClass().getSimpleName() + "." + method.getName();
LoggerFactory.getLogger(m).error("async-failure args={}", Arrays.toString(params), ex);
meters.counter("async.errors", "method", m).increment();
deadLetters.record(m, params, ex); // enable replay
};
}
}go deeper
Knows the default logs and that's usually not enough.
Can implement a custom handler with metrics and pick return types.
Adds context propagation, bounded executors, and retry/dead-letter thinking.
Defines the end-to-end async-failure policy (observability + recoverability + conventions + monitoring) and enforces it across teams.
## Framing The default behavior (void → `SimpleAsyncUncaughtExceptionHandler` logs ERROR; futures capture their own error) is fine for a demo but leaves two production gaps: **observability** (log-grepping doesn't scale, and unobserved futures log nothing) and **recoverability** (nothing retries or replays a failed side effect). A principal-level answer builds an explicit policy. ## 1. A real `AsyncUncaughtExceptionHandler` (for void tasks) Register it via `AsyncConfigurer.getAsyncUncaughtExceptionHandler()`. It should: - Log at ERROR with **structured context**: method name, args (scrubbed of secrets/PII), correlation/trace id. - **Emit a metric** (e.g. Micrometer `Counter` tagged by method) so alerting is on metrics, not logs. - For critical tasks, **alert** and/or **dead-letter**: persist a record (task type, payload, error, timestamp) to a table/queue for later inspection and replay. ```java @Override public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() { return (ex, method, params) -> { String m = method.getDeclaringClass().getSimpleName() + "." + method.getName(); log.error("async-failure method={} args={}", m, scrub(params), ex); meterRegistry.counter("async.errors", "method", m).increment(); deadLetterStore.record(m, params, ex); // for replay }; } ``` ## 2. Deliberate return-type policy - **void** for pure side effects the caller is indifferent to → handler covers failures. - **`CompletableFuture<T>`** when the result or failure must be consumed, aggregated, or fed back into the request. Always attach `exceptionally`/`handle`/`whenComplete` — an **unobserved `CompletableFuture` failure is swallowed with no handler and no log**, the single worst trap. Establish a team convention: "if you return a future, someone must consume it." ## 3. Context propagation Handlers and async bodies run on **worker threads**, so MDC, trace IDs, and `SecurityContext` are absent by default — making failure logs uncorrelatable. Fix with a `TaskDecorator` on the executor that snapshots and restores `MDC` and `SecurityContext` around the task. (Spring 6 / Micrometer `ContextPropagation` and `ContextSnapshot` formalize this.) ## 4. Executor and back-pressure Failure handling includes **rejection**: a `ThreadPoolTaskExecutor` with an unbounded queue hides overload; a bounded queue + explicit `RejectedExecutionHandler` (e.g. `CallerRunsPolicy` or a custom one that dead-letters) surfaces it. Name threads for diagnosability. ## 5. Retries / recovery Transient failures (webhook 503, SMTP blip) want retry, not just logging: `@Retryable` (Spring Retry) inside the async method, or re-enqueue from the dead-letter store with backoff. Ensure the task is **idempotent** so replays are safe. ## 6. Monitoring Alert on the `async.errors` metric and dead-letter depth — not on grep. Track per-method error rates. ## Anti-patterns to call out - Unobserved `CompletableFuture` → silent loss. - A custom handler that catches and returns without logging → destroys the only default signal. - Assuming the caller's `@Transactional` rolls back on async failure (it doesn't). - Relying on the handler for future-returning methods (it's void-only). - Blocking on `get()` right after the call, making async pointless. ## When to use what Fire-and-forget with a robust handler for non-critical side effects; futures + explicit recovery for result-bearing or business-critical async; retry/dead-letter for anything that interacts with flaky external systems.
- What's the single biggest silent-failure trap and how do you prevent it organizationally?An @Async method returning a CompletableFuture whose failure nobody observes — no handler fires and nothing is logged. Prevent it with a team convention (every returned future must be consumed), lint/review checks, and preferring void+handler when the caller genuinely won't consume the result.
- How do you make async-failure logs correlatable with the originating request?Add a TaskDecorator to the executor that snapshots MDC, trace/correlation IDs, and SecurityContext on the caller thread and restores them on the worker thread around task execution (Spring 6 / Micrometer ContextSnapshot formalizes this).
- Where would you add retry for a flaky webhook @Async task?Wrap the external call with Spring Retry (@Retryable + backoff) inside the async method for transient failures, and fall back to a dead-letter store for replay on exhaustion. Ensure idempotency so retries/replays are safe.
saying these in an interview costs you the question
- Relying solely on the default log-only handler in production
- Leaving CompletableFutures unobserved
- Writing a handler that catches and drops without logging or metrics
- Ignoring context propagation so failure logs have no trace id
- Unbounded executor queue that hides overload