How do you set a codebase-wide policy for which layers log exceptions and which re-raise?
answer
- State an invariant, not a preference
- Count the layers allowed to stop a failure
- Libraries have no boundary of their own
- Context should ride the exception
- Hooks are a net, not a design
basics
~10 sName a small, explicit set of boundary layers - request entry, worker loop, thread or process entry, CLI main - as the only places allowed to log an exception and stop. Everything else raises.
solid answer
~50 sWrite the invariant down: **one failure produces one error record**, emitted by the layer that stops it. Then enumerate the boundaries where failures legitimately stop — the entry point of a request or task, a worker loop that survives bad items, a thread or subprocess body, a callback crossing into foreign code, and `main()`. Those layers own `logging.exception`; every other layer raises, wrapping to add context if it has any. Library packages log nothing at ERROR and never call `sys.exit` — they raise their documented types. Make context travel with the failure rather than through extra log lines, so the boundary record is complete. Enforce with a review checklist and a lint rule for `except` blocks containing both a traceback log and a bare `raise`, and back it with `sys.excepthook` and `threading.excepthook` as a net for whatever escapes anyway — a net, never the design.
code
python · 13 linesimport logging
import sys
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("inventory.sync")
def last_resort(exc_type, exc, tb):
log.critical("unhandled, process is exiting", exc_info=(exc_type, exc, tb))
sys.excepthook = last_resort
raise TimeoutError("inventory sync aborted")go deeper
Understand that a codebase decides in advance which layers report errors, and that your function is usually not one of them. Raising is the normal thing to do; logging a traceback is the exception.
Be able to name the boundaries — request entry, worker loop, thread body, main — and explain why library code raises instead of logging. Know that context should be attached to the exception rather than logged separately.
Show how the rule survives contact with real teams: context carried on the exception, DEBUG breadcrumbs, deliberate log-and-rethrow only where a boundary loses the exception, and a lint rule so it does not decay.
Own the invariant across services and the tradeoffs behind it: duplication versus loss at process and queue boundaries, error-record counts feeding error budgets and alert routing, and why last-resort hooks are a net rather than an architecture.
## Start from the invariant, not the style guide A policy phrased as "do not over-log" produces arguments. A policy phrased as an invariant produces decisions: > **One failure, one error record, written by the layer that stops it.** Everything else follows. If a layer does not stop the failure, it does not write the error record. If nothing stops the failure, the missing thing is a boundary, not a log line. ## Enumerate the boundaries explicitly The policy is only usable if the list of layers permitted to log-and-stop is short, written down, and locatable in the code: - the entry point of a request or unit of work, where a failure becomes a response or a job status; - a worker or consumer loop that must survive one bad item and go to the next; - a thread body, a process entry point, or a callback invoked by code you do not own — anywhere an escaping exception would be swallowed or printed by a hook outside your control; - `main()`, which owns the exit status; - a scheduled task's outermost frame. Make these easy to see: a single decorator or a single helper used at all of them, so a reviewer can ask "is this one of the five?" and answer it by reading one import. Anything else that logs a traceback is a finding. ## The library rule is different, and stricter Code meant to be imported by other teams has *no* boundary of its own: it does not know whether its caller considers a failure fatal. So library code raises its documented exception types, does not log at ERROR, and never terminates the process. It may emit DEBUG for tracing, and it should leave handler configuration entirely to the application — a library that installs handlers or logs errors is making policy for someone else's process. ## Make context travel with the failure The strongest pressure against a single boundary log is real: the deep frame knows the row, the file, the tenant; the boundary knows none of it. If the policy has no answer, teams will re-add the intermediate logs. Two answers work. **Attach context to the exception.** Intermediate layers that know something raise a wrapping exception carrying it, chained to the original, so the one boundary record contains both the context and the root cause in a single traceback. **Propagate ambient context out of band.** Request or job identity that every log line should carry belongs somewhere the boundary reads automatically rather than in ad-hoc messages, so the single record is already attributable. With either in place, an intermediate log is no longer buying anything, and the rule holds under pressure. ## Enforcement Convention decays; make the machine carry some of it. - A lint rule over `except` blocks that contain both a traceback-bearing log call and a bare `raise` catches the exact anti-pattern mechanically, with an allowlist comment for the deliberate cases. - A review checklist item — *which layer stops this failure?* — turns a stylistic argument into a factual one. - Dashboards close the loop: if one incident produces four error events, the duplication is visible without reading code. ## The last-resort net `sys.excepthook` handles what reaches the top of the main thread; `threading.excepthook` (Python 3.8+) covers threads, whose exceptions `sys.excepthook` never sees. Installing both, logging at CRITICAL, is worth doing — it converts an unstructured stderr dump into a record your pipeline actually collects. But it is a net for gaps, not a design: a hook cannot produce a response, cannot mark a job failed, and cannot decide an exit status. If failures routinely reach it, boundaries are missing. ## The tradeoffs a lead actually owns - **Duplication versus loss.** A strict single-record rule risks a failure crossing an unnoticed boundary — a queue, a subprocess, a foreign callback — and disappearing. Audit those transitions specifically; they are the one place an extra deliberate log is correct. - **Debuggability versus volume.** Teams that have been burned by a silent failure want to log everywhere. Buy that back with cheap DEBUG breadcrumbs and rich exception context rather than with ERROR records. - **Metric integrity.** Duplicate error events inflate error rates and distort error-budget arithmetic; a service whose deep layers log-and-rethrow can look several times less reliable than it is, or, once people learn to divide, mask a real regression. - **Consistency across services.** The value of the invariant is cross-service: on-call reading two services' logs should be able to assume the same relationship between records and failures. That is why this is a policy and not a per-repo preference.
- How do you answer the team that says logging deep gives them the context the boundary lacks?Agree with the need and move it off the log. Intermediate layers raise a wrapping exception carrying the identifying detail, chained to the original, so the single boundary record shows context and cause together; ambient identity such as a request or job id is propagated so every record already carries it. If they still want breadcrumbs, DEBUG without a traceback is cheap and off in production.
- Does sys.excepthook cover exceptions raised in worker threads?No. Exceptions escaping a thread's run go to threading.excepthook, added in Python 3.8; sys.excepthook only sees the main thread. Install both if you want a genuine net. Neither is a substitute for a handler inside the worker loop, because a hook cannot mark a task failed, produce a response or choose an exit status — it can only record what already went wrong.
- How would you detect violations of this policy without reading every diff?Two loops. Statically, a lint rule flagging except blocks that both log with a traceback and bare-raise, with an allowlist comment for deliberate boundary crossings. Operationally, compare error-record counts against incident counts: an outage that produces four records per failure shows the duplication without anyone reading code, and the same signal catches the opposite defect when a failure produces none.
saying these in an interview costs you the question
- Every layer logs defensively in case the next does not
- Library code logging at ERROR or calling sys.exit
- Treating sys.excepthook as the error-handling design
- Assuming sys.excepthook covers worker threads
- Leaving the policy to convention with no enforcement
- Ignoring what duplicate records do to error-rate metrics