skip to content

Why does calling logger.exception() and then re-raising at every layer report one failure many times?

level: middleimportance: must knowfreq 50%

answer

  1. Re-raising consumes nothing
  2. Count the reports, not the failures
  3. Each handler on the way out writes one
  4. Handle it or propagate it, never both
  5. Log where the failure actually stops

basics

~20 s

A re-raised exception keeps travelling with its full traceback, so every layer that logs it emits a complete report of the same failure. One error becomes N log records, N alerts and N counted errors. Log it once, where it stops.

solid answer

~50 s

Re-raising does not consume anything: the exception keeps its traceback and continues upward, so each `except: logger.exception(...); raise` on the way out writes another full report of one failure. The log then shows four ERROR records and four tracebacks for a single timeout, error-rate metrics count it four times, and alerting fires repeatedly. The rule is **handle it or propagate it, never both**: if you can recover, catch, log once and return a defined result; if you cannot, let it propagate — the exception is already carrying everything the top-level handler needs. Add context on the way up by raising a wrapping exception rather than by logging. The legitimate exceptions to the rule are the layers where propagation actually ends: a request handler at the top, a worker or thread body, a callback that crosses into foreign code.

code

python · 26 lines
python
import logging

logging.basicConfig(level=logging.INFO)
log = logging.getLogger("inventory.sync")


def fetch():
    raise TimeoutError("upstream inventory feed timed out")


def middle():
    try:
        fetch()
    except TimeoutError:
        log.exception("middle: fetch failed")   # report 1
        raise                                   # same exception continues


def top():
    try:
        middle()
    except TimeoutError:
        log.exception("top: request failed")    # report 2, same failure


top()

go deeper

for a junior

Remember that a re-raised exception keeps its traceback and keeps going, so logging it on the way out just repeats the same report. Be able to say where the single log call belongs.

for a middle

Explain the handle-or-propagate rule and what each branch owes: recovery logs once because the failure becomes invisible, propagation stays silent because something above will report it. Know why catch-log-return-None is the worst option.

for a senior

Show how you add context on the way up by raising rather than logging, and name the boundaries where a log-and-rethrow is deliberate. Talk about what duplicate records do to error rates and alert fatigue.

for a principal

Own the invariant across services: one failure produces one error event. Decide which layers are designated boundaries, how that is enforced in review, and how duplicate reporting distorts error budgets and on-call routing.

## What re-raising does and does not cost `raise` inside an `except` block re-raises the exception being handled, unchanged: same object, same `__traceback__`, with the current frame appended as it unwinds. Nothing is consumed, nothing is reset. That is precisely why *log-and-rethrow* duplicates: each handler on the unwind path that calls `logger.exception()` writes a full report — message plus traceback — of the identical failure, and then hands the failure on so the next layer can do it again. A four-layer stack turns one intermittent timeout into four ERROR records with four nearly identical tracebacks (each slightly longer than the last), four increments of whatever counts errors, and four inputs to the alert pipeline. On-call then has to work out whether the service saw one problem or four. Deduplicating that after the fact is far harder than not creating it. ## The rule **Every `except` block owes exactly one of two things: recovery or propagation.** - **Recovery** — you can produce a correct result without the failed operation: fall back to a cache, skip the record, return a default. This is where a log belongs, because the failure is about to become invisible: nobody upstream will ever see it. Log once, at a level that matches how much you care, and carry on. - **Propagation** — you cannot fix it. Then do not log; let it go. The exception is a *complete report already*, and it is heading toward a layer whose job is to report it. Catching purely to log is taking custody of a failure you have no intention of handling. The pathological middle ground is *catch, log, return `None`* — the caller sees a normal return, proceeds on missing data, and fails later somewhere unrelated. The log record exists but nothing acted on it. That is worse than the duplication problem, because it converts a loud failure into a quiet corruption. ## Where the logging call legitimately lives Exactly at the places where propagation *ends*, because there is nothing above to report it: - the top of a request or unit of work, where the failure is turned into a response or a job status; - a worker loop that must survive one bad item and continue to the next; - a thread body or a callback invoked by foreign code, where an escaping exception would be swallowed by the framework or printed by a hook you do not control; - `main()` in a CLI, before choosing an exit status. One of those, per failure. Everything between the `raise` and that boundary stays silent. ## Adding context without adding a log line The usual defence of log-and-rethrow is "the middle layer knows things the top does not — which file, which row, which tenant." That is a real need, and logging is the wrong tool for it, because it splits one failure across two records that nothing joins. Two better moves: - **Raise with the context attached.** Wrap the low-level error in an exception carrying the identifying detail and chain it, so the top-level report shows both the context and the original cause in one traceback. - **Propagate context out of band.** Put the request or job identity somewhere the boundary log will read anyway, so one record carries it. If you truly want breadcrumbs at intermediate layers, log them at **DEBUG** without `exc_info`: cheap, off in production, and unmistakably not an error report. ## The narrow, deliberate exception Sometimes a layer must log *and* re-raise: it is about to cross a boundary where the exception will be mangled or lost — pushed through a queue, marshalled to another process, handed to code that catches broadly. There the extra record is the point, not an accident. Say so in a comment, and prefer a distinct level or message so the duplicate is recognisable as intentional rather than looking like a second failure. ## Reviewing for it The pattern is easy to grep for: an `except` clause whose body contains both a logging call with a traceback and a bare `raise`. When you find one, ask which layer *stops* this failure. If the answer is "a handler above", delete the log. If the answer is "nothing — it reaches the top and dies", the fix is to add the missing boundary handler, not to log at every floor on the way down.

  • The middle layer knows which file failed and the top layer does not. How do you keep that detail without logging twice?
    Attach it to the exception instead of to the log: raise a wrapping exception whose message carries the file, chained from the original, and let it propagate. The single boundary log then renders the context and the original cause in one traceback. If you only want a breadcrumb, log at DEBUG without exc_info — it is off in production and reads as a trace, not an error report.
  • When is catching an exception, logging it, and re-raising actually correct?
    When the layer is about to cross a boundary that loses the exception: pushing work through a queue, marshalling to another process, or handing control to foreign code that catches broadly. There the extra record is the last chance to capture the traceback, so it is deliberate rather than accidental. Mark it as intentional in the code and make its message distinguishable from the boundary handler's.
  • What is wrong with catching an exception, logging it, and returning None?
    The caller cannot tell failure from a legitimately empty result, so it continues on missing data and fails later somewhere unrelated — usually with a traceback that points nowhere near the cause. The log record exists but nothing acted on it. Either recover to a genuinely valid result, or let the exception propagate so a layer that can decide gets the chance.

saying these in an interview costs you the question

  • Believing re-raising clears or consumes the traceback
  • Log-and-rethrow in every layer for safety
  • Catching an exception only in order to log it
  • Catching, logging, and returning None to callers
  • Treating each duplicate log record as a separate failure
  • Adding context by logging instead of by raising

context