skip to content

What is SLF4J's Mapped Diagnostic Context (MDC), and why does putting a correlation identifier into it often produce missing or wrong values in a service that uses thread pools and asynchronous work?

level: middleimportance: should knowfreq 55%

answer

  1. per-thread key-value map printed by the layout
  2. set at boundary, clear in finally
  3. pooled thread reuse = stale wrong context
  4. hand-off = empty map, capture and restore
  5. inheritable thread-locals do not help pools

basics

~20 s

MDC is a per-thread key-value map that the layout can print on every line, so you set a correlation id once instead of passing it to every log call. It is thread-local: work handed to another thread does not inherit it, and a pooled thread keeps stale entries unless cleared.

solid answer

~60 s

MDC is a small map attached to the current thread. You put entries at the boundary (request id, tenant, user), the backend's layout or structured encoder emits them on every event from that thread, and you remove them when the unit of work ends. SLF4J only defines the API; the bound backend supplies the adapter and decides how the map is stored and whether child threads inherit it. Two failure modes follow from the thread-local storage. First, **leakage**: pooled threads are reused, so an entry not removed in a finally block reappears on the next unrelated request, attributing lines to the wrong trace. Always clear in a finally, or clear the whole map at the pool boundary. Second, **loss**: when you submit work to an executor, publish an event, or continue on a callback thread, the new thread has its own empty map. You must capture the copy of the context map before handing off and restore it on the executing thread — usually via a wrapping executor or task decorator so application code never does it by hand.

code

text · 10 lines
text
boundary:
  put("requestId", id)
  try { handle() } finally { clear() }

hand-off:
  captured = copyOfContextMap()          // on the submitting thread
  submit(task {
     setContextMap(captured)             // on the executing thread
     try { run() } finally { clear() }
  })

go deeper

for a junior

Say MDC is a per-thread map you fill at the start of a request so every log line automatically shows the request id, and that it must be cleared afterwards.

for a middle

Explain both failure modes concretely: stale context on reused pool threads, and empty context after a hand-off, plus the copy/restore fix.

for a senior

Argue for centralising the lifecycle in a boundary component and a task-decorating executor, and cover the interaction with a tracing library and with asynchronous appenders.

for a principal

Set the platform convention: which keys exist, who owns them, that they are identifiers rather than payload, and how propagation is guaranteed by shared infrastructure instead of by developer discipline.

## What MDC is for Log lines are useful in isolation only for a single-threaded toy. In a service, the question is always "show me everything that happened while handling request X". MDC exists so that the identifying context is attached once, at the boundary, and then printed automatically on every subsequent line rather than threaded through every method signature and every log call. The API is deliberately tiny: put a key and value, get a value, remove a key, clear everything, and take or set a copy of the whole map. The layout side is the backend's job — a pattern layout references the keys by name, and a structured encoder emits them as fields. SLF4J itself only defines the facade; the bound backend provides the adapter implementation, which is why behaviour details (notably whether new threads inherit the map) vary by backend and version. ## Why it is thread-local, and what that buys Binding context to the thread means no parameter has to be passed and no logging call has to change. In the classic one-thread-per-request model this is close to free and close to correct: the thread *is* the unit of work. Everything that goes wrong with MDC goes wrong because that assumption breaks. ## Failure mode 1: leakage across pooled requests Application threads are pooled. If you put a request id and the request ends without removing it — an early return, a thrown exception, a filter that only cleans up on the happy path — the entry stays on the thread. The next request handled by that thread inherits it, and now log lines carry another request's identifier. This is worse than having no context: it produces confident, wrong attribution, and it also produces genuine privacy incidents when the leaked value is a user or tenant identifier. The defences are ordered. Best is a single boundary component that sets the context and clears it in a `finally`, so no application code touches MDC at all. Next best is clearing the entire map (not just your keys) at the start of each unit of work, which is cheap insurance against anything a library left behind. Removing individual keys by hand at many call sites is the fragile option, because the set of keys drifts. ## Failure mode 2: loss across a thread hand-off Any hand-off creates a thread with an empty context: submitting to an executor, a parallel stream, a scheduled task, a reactive operator that switches schedulers, a completion callback from an HTTP or database client, a message-listener thread. The unit of work continues but the diagnostic context does not, so the tail of the request logs with no identifier at all — often exactly the part you needed, because the asynchronous part is where the failure was. The fix is capture-and-restore: on the submitting thread take a copy of the context map, and on the executing thread set it before the task body and clear it afterwards. Doing this in application code is unreliable because it must be remembered at every hand-off. The maintainable form is an executor wrapper or task decorator applied once, plus wrappers for any framework-specific dispatch points. Inheritable thread-local storage is not a general answer: it copies only at thread *creation*, which is precisely when a pool does not create threads, so it appears to work in a test and fails under load. ## Interaction with tracing and structured logging MDC is where a trace and span identifier is usually surfaced to the log line, so a log search can pivot to a trace. Tracing libraries typically own propagation across threads themselves and mirror their context into MDC; when they do, do not also propagate the same keys by hand, or the two mechanisms will disagree at the edges. Keep MDC keys few, low-cardinality in *name* (the values are naturally high-cardinality) and documented, because the layout must reference them by name and unknown keys silently render as empty. ## Cost and limits MDC values are strings; store an identifier, not a serialised object. The map is copied at hand-off points, so keeping it small matters when hand-offs are frequent. And because the map is read at format time, an asynchronous appender must snapshot the context with the event — a good backend does this automatically, but it is another reason not to mutate MDC entries after logging.

  • Why is an inheritable thread-local not a general solution for propagating MDC into a thread pool?
    Inheritable context is copied from the parent at the moment a thread is created. A pool creates its threads once, typically while handling some arbitrary early task, and then reuses them for everything afterwards, so later submissions inherit nothing — or worse, inherit the context that happened to exist when the pool grew. It looks correct in a small test where each task spawns a fresh thread and fails in production. Propagation must happen per task, at submit time, not per thread at creation time.
  • A log line shows a request id belonging to a completely different user. How did that happen and how do you prevent it?
    A previous unit of work on that pooled thread put the value into MDC and never removed it, most often because cleanup was on the success path rather than in a finally block, or because a library set a key nobody knew about. Prevention is to own the lifecycle in one boundary component that clears the whole context map at the start and again in a finally, rather than removing individual keys at many call sites. Treat it as a correctness and privacy bug, not cosmetic noise.
  • How does MDC relate to the fields emitted by a tracing library?
    Tracing libraries maintain their own context and propagate it across threads and processes; most of them mirror the trace and span identifiers into MDC so that log lines are joinable with traces. In that arrangement MDC is a projection for the log layout, not the source of truth. Do not also propagate those same keys yourself, or the two mechanisms can disagree at hand-off boundaries and produce lines stamped with a stale trace id.

MDC is a wristband put on at the door of a venue: everything the wearer does is stamped with it automatically. It stays on the person, not the task, so if someone else picks the task up they carry no band — and if the band is never cut off, the next visitor gets handed a used one.

saying these in an interview costs you the question

  • Believing MDC is inherited automatically by tasks submitted to an executor.
  • Clearing MDC only on the success path instead of in a finally block.
  • Storing large objects or secrets as MDC values rather than short identifiers.
  • Assuming SLF4J itself decides how MDC is stored or printed, rather than the bound backend and its layout.
  • Reaching for inheritable thread-locals as the fix for pooled executors.

context