skip to content

You wrap a Logback file appender in AsyncAppender to cut latency. Explain the queue model, what queueSize, discardingThreshold, neverBlock and includeCallerData actually do, and which log lines you can silently lose.

level: seniorimportance: should knowfreq 45%

answer

  1. bounded queue 256, one worker, decorator over exactly one appender-ref
  2. discardingThreshold default = size/5: drops INFO and below at 80% full
  3. neverBlock=false blocks the app thread; true drops
  4. caller data must be captured on the calling thread — default off, prints ?
  5. shutdown drains only up to maxFlushTime (1000 ms)

basics

~20 s

AsyncAppender puts events on a bounded blocking queue (default 256) drained by one worker thread. When the queue is 80% full it silently drops TRACE/DEBUG/INFO events by default; when full it blocks the caller unless neverBlock=true, which drops instead. Caller data is not captured unless enabled.

solid answer

~50 s

`AsyncAppender` is a decorator: the calling thread enqueues the event onto a bounded `BlockingQueue` (`queueSize`, default 256) and returns; a single worker thread dequeues and forwards to the wrapped appender. Four knobs matter. **`discardingThreshold`** defaults to `queueSize/5` — once remaining capacity falls below it, events of level INFO and below are dropped silently, so warnings and errors survive a burst but your INFO trail vanishes exactly during the incident. Setting it to `0` disables dropping. **`neverBlock`** decides what happens when the queue is genuinely full: default `false` means the application thread *blocks* on `put`, converting a slow disk or a stalled network appender into application latency; `true` swaps that for silent loss. **`includeCallerData`** defaults to false because caller data must be captured on the *calling* thread — the worker's stack is useless — so `%line`/`%method` print as `?` unless enabled, and enabling it reintroduces the stack-walk cost you were avoiding. On shutdown, undrained events are lost unless the context is stopped cleanly within `maxFlushTime`.

code

text · 5 lines
text
remaining > 51        -> enqueue everything
remaining <= 51       -> INFO/DEBUG/TRACE dropped silently; WARN/ERROR enqueued
remaining == 0        -> neverBlock=false: caller blocks until a slot frees
                         neverBlock=true : event dropped, caller continues
JVM exit w/o context stop -> whatever is still queued is lost

go deeper

for a junior

Know that AsyncAppender hands the event to a background thread via a bounded queue, so logging returns faster but lines can be lost.

for a middle

Be able to name queueSize (256), the queueSize/5 discarding threshold that drops INFO and below, neverBlock, and why %line stops working.

for a senior

Frame it as an availability-versus-observability policy: pick blocking or dropping deliberately, size the queue against heap and burst shape, disable discarding when the trail matters, and make sure shutdown drains within maxFlushTime.

for a principal

Argue when async logging is the wrong tool at all — durable or audit streams need a real guarantee, not a queue — and set an org-wide default (neverBlock plus drop alerting for request-path services) so failure modes are uniform and diagnosable across the fleet.

## What it actually decouples Synchronous appenders write on the caller's thread and typically flush per event, so the application pays a syscall and, under contention, an appender lock. `AsyncAppender` moves the write off that thread. What it does **not** do is make logging free: enqueueing still costs, the queue is shared, and every guarantee you had about a line being on disk when the method returned is gone. Note also that `AsyncAppender` writes nothing itself. It must reference exactly one wrapped appender via a single `<appender-ref>`, and that appender does the real I/O. Pointing it at several appenders is a configuration mistake people make and then wonder why only one destination receives events. ## The queue A single bounded blocking queue of `queueSize` events, drained by one worker thread. Defaults are conservative: 256 entries. Two implications: - 256 is small for a bursty service. A burst larger than the queue hits the back-pressure or drop path within milliseconds. - One worker means ordering is preserved but throughput is bounded by the single downstream appender. Sizing is a memory decision: each queued logging event retains the message, argument references and the copied MDC map, so a very large queue on a high-cardinality MDC can hold real heap — and, worse, keeps references to argument objects alive, delaying their collection. ## discardingThreshold: the silent, level-biased drop When the queue's remaining capacity falls below `discardingThreshold`, `AsyncAppender` discards incoming events at level INFO, DEBUG and TRACE, keeping WARN and ERROR. The default threshold is `queueSize / 5`, i.e. dropping starts at 80% full — with the default 256 that is a threshold of 51, so once fewer than 51 slots remain, INFO and below stop being enqueued. The design intent is good — preserve the important lines under pressure — but the operational consequence is precisely what surprises teams: during the incident, when volume spikes, the contextual INFO lines that explain *why* the errors happened are the ones deleted, and nothing in the output says so. If you need a complete trail, set `discardingThreshold=0` and accept that you now face the full-queue decision instead. ## neverBlock: back-pressure versus loss With the queue full and no discarding available: - `neverBlock=false` (default) → the enqueue blocks the application thread until space appears. Your logging subsystem is now on the critical path, and a stalled destination (an NFS mount, a full disk, a TCP appender to a down collector) becomes an application-wide latency or thread-pool exhaustion event. - `neverBlock=true` → the non-blocking offer fails, the event is dropped, the caller continues. Availability is protected; observability is not. There is no third option inside `AsyncAppender`. Choosing is a policy call, and the honest framing is: *is a missing log line or an added tail-latency spike the worse outcome for this service?* For request-path services the usual answer is `neverBlock=true` plus alerting on drops; for audit or compliance streams the answer is blocking, or better, not using an async appender at all and writing durably instead. ## includeCallerData Caller information (class, method, file, line) is derived by walking the current thread's stack. Once the event is on the queue, the worker's stack has no relation to the origin, so the data must be captured **eagerly, on the calling thread**, at the moment of enqueue. That is exactly the expensive thing you moved to async to avoid, which is why the default is `false` — and why patterns containing `%line` or `%method` render `?` behind an `AsyncAppender` until someone sets `includeCallerData=true`. ## Shutdown Events sitting in the queue when the JVM exits are lost unless the logger context is stopped, which stops the appender and waits up to `maxFlushTime` (default 1000 ms) for the worker to drain. In containers, a hard kill after a short grace period, or an app that never stops the context, loses the tail — typically the shutdown or crash lines you most wanted. ## When to reach for it Async helps when the wrapped appender is genuinely slow or blocking (network appenders, remote syslog, heavily formatted output) or when tail latency is the target. For a plain local file appender, turning off per-event flushing often recovers most of the throughput with a far simpler failure model. Measure before adopting a component whose failure mode is *silence*. Worth knowing for contrast: Log4j2 takes a different design for the same goal, with an asynchronous logger built on a ring buffer (LMAX Disruptor) and configurable wait strategies, rather than Logback's single bounded queue plus level-biased discard. The trade-off space is the same — bounded memory, back-pressure or loss — but the knobs and defaults do not transfer between the two.

  • After adding AsyncAppender, %line and %method print as ?. Why, and what is the fix's cost?
    Caller data comes from walking the calling thread's stack, and by the time the worker thread formats the event that stack is gone, so Logback records nothing unless it captured it at enqueue time. Setting includeCallerData=true makes the capture happen on the application thread, restoring the fields. The cost is a stack walk per event on the hot path, which is one of the largest single costs in logging and undermines much of the reason for going async.
  • Under load your INFO lines disappear but errors still arrive. What is happening and how do you confirm it?
    That level-biased pattern is the AsyncAppender discarding threshold: once remaining capacity drops below queueSize/5 it silently drops INFO and below while keeping WARN and ERROR. Confirm it by correlating the gaps with load — the loss is bursty and stops when traffic falls — and by setting discardingThreshold to 0 in a test run and seeing the INFO lines return (now paying back-pressure or full-queue drops instead). Logback's own status listener can also be enabled to surface appender-level warnings.
  • Would you use AsyncAppender for an audit log?
    Generally no. AsyncAppender's contract is that an event may be dropped or lost at shutdown, and neither discarding nor neverBlock gives you a durability guarantee, so an audit trail behind it is a trail with holes you cannot detect. If audit records matter, write them synchronously to a durable sink, or push them through a store designed for the guarantee rather than through the logging framework's throughput optimisation.

It is a conveyor belt with a fixed number of slots: when it backs up you either hold the worker at the belt (blocking) or throw the parcel away (neverBlock), and by default the cheap parcels get thrown away first.

saying these in an interview costs you the question

  • Believing AsyncAppender never loses events — it drops INFO and below at 80% full by default, and drops the whole queue on an unclean shutdown.
  • Thinking neverBlock=true is strictly safer; it converts back-pressure into silent loss, which is the right trade only for some services.
  • Assuming caller data (%line, %method) works the same behind an async appender — it must be captured on the calling thread and is off by default.
  • Treating async as a fix for a slow or unavailable destination; it only buys queueSize events of slack before blocking or dropping.
  • Pointing one AsyncAppender at several appender-refs and expecting fan-out — it wraps exactly one appender.

context