skip to content

How do you turn a service's log level up at runtime without a redeploy, and why per-logger rather than globally?

level: middleimportance: should knowfreq 44%

answer

  1. Detail on demand, not on always
  2. The change must reach live objects
  3. Cached effective levels need recomputing
  4. One package, one replica, time-boxed

basics

~20 s

A runtime change has to mutate the live logger objects, not just a configuration file, and invalidate the effective level cached on every descendant logger. Scope it per-logger so one package gets detail while everything else stays at INFO.

solid answer

~40 s

Production normally runs at INFO, so the real question is how to get more detail without a restart — a restart discards the process state you were investigating. A runtime change has to reach the live logger objects: resolve the named logger, set an explicit level on it, and recompute the effective level cached on every descendant, because the hot-path level check reads a cached field rather than walking the hierarchy each call. Editing the deployed configuration file alone does nothing, since the process built its loggers from it at startup. Scope the change per logger — `inventory.reconcile.humidity`, not the root — because global DEBUG on a busy service can multiply output enormously while it is already unhealthy. And remember the change is process-local and dies on restart, so time-box it and revert deliberately.

code

pseudocode · 9 lines
pseudocode
setLevel("inventory.reconcile.humidity", DEBUG):
    logger := registry.getOrCreate("inventory.reconcile.humidity")
    logger.configuredLevel := DEBUG

    # the hot path reads effectiveLevel as a field, so every
    # inheriting descendant must be recomputed or the change is invisible
    for child in registry.descendantsOf(logger):
        if child.configuredLevel == INHERIT:
            child.effectiveLevel := resolveFromNearestConfiguredAncestor(child)

go deeper

for a junior

Know that production normally runs at INFO and that the level can be changed while a service is running. Be able to say why restarting a process to get more detail is usually the wrong move during an incident.

for a middle

Explain the mechanics: loggers form a hierarchy by name, a logger without an explicit level inherits its nearest configured ancestor's, and implementations cache that resolved level, so a runtime change must recompute the cache on descendants.

for a senior

Demonstrate incident judgement. Target one logger and ideally one replica, time-box the change, watch output rate and appender pressure while it is on, and revert deliberately so the fleet does not drift into undocumented per-process configuration.

for a principal

Own it as a platform capability: every service exposes the same control surface, changes are audited and expire on their own, and the shared collection tier carries enough headroom that one team turning on detail cannot degrade everyone else's telemetry.

## Why production defaults to INFO INFO is the level at which a service tells the story of what it did without narrating how: state transitions, configuration resolved at startup, lifecycle changes, a batch finishing, a notable outcome. It is enough for the next reader to reconstruct the shape of what happened, at a rate that stays readable when 41 services write into the same place. DEBUG as a standing default fails on three counts at once, and only one of them is money: - **Per-request cost.** Every emitted line costs an argument-formatting pass and a handful of allocations. Where code concatenates eagerly instead of passing a template plus its arguments, that cost is paid even for lines the level check then discards. - **Backpressure.** Appenders have bounded queues. When one fills, an implementation either drops — losing the lines you turned on — or blocks, which moves log-writing cost onto the request thread. That is how a logging change becomes a latency incident. - **Signal.** A reader hunting the one line that explains an incident does worse, not better, with fifty times as much material. So the useful arrangement is not "verbose or quiet" but **quiet by default, verbose on demand, narrowly** — which only works if the level can change while the process is running. ## What a runtime change has to reach Editing the configuration file that shipped with the deployment changes nothing on its own: the process read it once, at startup, and built objects from it. A working runtime change has to do three things inside the live process. 1. **Resolve the named logger.** Loggers form a hierarchy by dotted name, so `inventory.reconcile.humidity` is a descendant of `inventory.reconcile`, which descends from `inventory` and ultimately from the root logger. 2. **Set an explicit level on that node.** A logger without one has no level of its own; it inherits from the nearest configured ancestor. 3. **Invalidate the effective level cached on every descendant.** This is the step people forget. Because a level check runs on every single log call, implementations do not walk the hierarchy each time — they cache a resolved effective level on each logger and read a field. Changing an ancestor without recomputing its descendants' caches leaves the change looking ignored. How the instruction gets in varies: a control surface exposed by the application, a signal the process handles, a configuration file the process watches, or a central configuration service the process subscribes to. The transport is a detail; the three steps above have to happen regardless. ## Why per-logger rather than global | Scope | What you get | What it costs | |---|---|---| | Redeploy with new configuration | Durable, reviewable, in the repository | Minutes at best, and a restart discards the process state you were investigating | | Global runtime change to DEBUG | Everything, immediately | Every code path narrating at once; the flood can outrun the collection tier and lose the lines you wanted | | Per-logger runtime change | Detail from the suspect code only | You need a hypothesis about where to look | The middle row is the trap. On a busy service, a root-level change to DEBUG can multiply output by two orders of magnitude within a second, and it does so while the service is already unhealthy. Per-logger granularity turns a blunt and dangerous instrument into an ordinary diagnostic step: one package at DEBUG produces a stream you can actually read, from exactly the code under suspicion, while the rest of that service and the other forty stay at INFO. ## Operating the change during an incident 1. **Form a hypothesis first.** "Which logger?" has an answer only if you have a suspect. Otherwise you are switching on DEBUG to browse, which is how the flood happens. 2. **Narrow the blast radius.** One logger, and where the control surface allows it, one replica. A single instance behind a load balancer still sees a share of traffic, and a reproducible case can be aimed at it deliberately. 3. **Watch what the change costs while it is on** — output rate, plus any queue-depth or dropped-line signal the runtime exposes. If the volume is unsustainable you want to know in seconds, not after the incident. 4. **Time-box it.** Decide up front when it reverts, before the detail arrives and attention moves elsewhere. 5. **Revert deliberately** and record what the extra detail actually told you. ## What a runtime change is not It is not configuration. It lives in one process's memory, it is not in the repository, and it disappears the moment that process restarts — which on an autoscaled or rolling-deployed platform may happen minutes later, silently. That property is a feature during an incident, because nothing is left behind by accident, and a hazard afterwards, because a fleet where three instances are verbose and thirty-eight are not has no record of why. If a package genuinely needs more detail permanently, that belongs in committed configuration, reviewed like any other change. The runtime control exists for the twenty minutes in which you need to see something INFO does not say.

  • The level change applied cleanly but the extra detail never appeared in the aggregated view. What went wrong?
    Most often the change is process-local and the traffic you are watching is served elsewhere. A control surface applied through a load balancer reaches one instance, and the requests you are chasing may land on any of the others. Either target that instance deliberately and reproduce against it, or push the change through a mechanism that reaches every replica and accept the volume that implies.
  • Why not simply leave the elevated level in place after the incident?
    Because it is invisible state. A runtime change lives in process memory, is not in the repository, and vanishes at the next restart, so the fleet quietly ends up in two configurations with nothing recording which. Set an expiry, revert deliberately, and if a package genuinely needs more detail permanently, change the committed configuration so the next deploy carries it.
  • What is the risk of switching the root logger to DEBUG on a busy service mid-incident?
    You can turn a degradation into an outage. Verbose logging adds formatting and allocation to every request path, and if the appender blocks when its queue fills, that cost lands on request threads. Downstream, a sudden jump in volume can outrun the collection tier and lose the very lines you turned on, along with other services' lines sharing that tier.

It is a dimmer on one lamp, not the mains switch for the whole building.

saying these in an interview costs you the question

  • Says you must redeploy to change a log level
  • Restarts the process to pick up logging config mid-incident
  • Switches the root logger to DEBUG as a first move
  • Assumes editing the config file changes a running process
  • Forgets the runtime change is process-local and dies on restart
  • Leaves the elevated level on indefinitely afterwards