skip to content

Failures at API Boundaries

The decisions a failure forces at the edge of a module: which foreign errors to translate into your own types, what your package promises to raise, and which calls a caller may safely repeat.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

21

What does logging.exception() record that logging.error() alone does not?

level: juniorimportance: must knowfreq 58%

answer

  1. One of them attaches more than a message
  2. Think about what a level shortcut hides
  3. It only means something inside except
  4. exc_info=True, done for you, at ERROR
  5. stack_info answers a different question

basics

~10 s

logging.exception() logs at ERROR level and attaches the traceback of the exception currently being handled, because it passes exc_info=True for you. logging.error() logs only the message unless you pass exc_info=True yourself.

solid answer

~40 s

`logging.exception(msg)` is exactly `logging.error(msg, exc_info=True)`: same ERROR level, plus the type, value and traceback of the exception being handled right now. It is only meaningful **inside an `except` block** — called anywhere else it appends the line `NoneType: None`, because there is no active exception to describe. `exc_info` also accepts an exception instance or a `(type, value, traceback)` tuple, so you can attach a traceback to any level: `log.warning(msg, exc_info=exc)` or `log.critical(msg, exc_info=True)`. A separate keyword, `stack_info=True`, attaches the call stack that reached the logging call itself, with no exception involved — that answers "who called this?", not "what blew up?". Put the context the traceback lacks (which record, which batch, which URL) in the message; the traceback already carries the exception text.

code

python · 11 lines
python
import logging

logging.basicConfig(level=logging.INFO)
log = logging.getLogger("inventory.sync")

try:
    int("6,800")
except ValueError:
    log.error("row count unparsable")                   # message only
    log.exception("row count unparsable")               # message + traceback
    log.error("row count unparsable", exc_info=True)    # identical to the line above

go deeper

for a junior

Recall the identity: logging.exception() is logging.error() plus exc_info=True, and it belongs inside an except block. Be ready to say what extra text lands in the log because of it.

for a middle

Explain what exc_info resolves to, that it also accepts an exception instance or a triple, and that stack_info attaches the caller's stack instead of a traceback. Know why a %-style message beats an f-string.

for a senior

Show judgement about level: a swallowed, retriable failure is a WARNING with a traceback, not an ERROR. Talk about putting identifying context in the message and letting the traceback carry the rest.

for a principal

Own the log-shape contract: one exception should produce one structured record with the traceback as a field, not text glued into a message, so records can be grouped, redacted and routed consistently across services.

## The one-line identity `Logger.exception(msg, *args)` is a convenience wrapper. Its entire behaviour is: log at `ERROR` level with `exc_info=True`. Nothing else. Writing `log.exception("sync failed")` and `log.error("sync failed", exc_info=True)` produces byte-identical output. The method exists because logging a caught exception without its traceback is the single most common logging mistake, and a dedicated verb makes the right thing shorter than the wrong thing. ## What `exc_info` actually resolves to When `exc_info` is truthy, the logging machinery captures the *currently handled* exception — the same thing `sys.exc_info()` returns — and stores the `(type, value, traceback)` triple on the log record. A formatter later renders it as the familiar `Traceback (most recent call last): ...` block underneath the message. "Currently handled" is the crucial qualifier. Inside an `except` block there is one; outside, there is not. Call `log.exception("oops")` at module level and you get the message followed by the literal line `NoneType: None` — a puzzling artefact that always means *you logged a traceback where no exception was live*. The same applies after the `except` block has finished: Python clears the exception state on the way out, so a deferred `log.exception()` records nothing useful. `exc_info` is not limited to `True`. It also accepts: - an **exception instance** — `log.error("deferred", exc_info=exc)`, useful when you stashed the exception and log it later, or when you are reporting an exception you never caught in this frame (one pulled off a queue, or a `Future`'s stored exception); - an explicit **`(type, value, traceback)` tuple** — what a top-level hook such as `sys.excepthook` is handed. Because `exc_info` is an ordinary keyword on every level method, the traceback is not welded to ERROR. A retriable failure that you intend to swallow is honestly a `WARNING` *with* a traceback; a failure that is about to kill the process is `CRITICAL` with one. `logging.exception()` is only the ERROR-level shorthand, and reaching for it purely to get the traceback is how logs end up full of ERROR entries for things that were handled fine. ## `stack_info` is a different question `stack_info=True` is often confused with `exc_info` because both append a multi-line block. They answer different questions: - `exc_info` → *where did the exception come from?* It renders the traceback of a raised exception, unwinding from the `raise` site. - `stack_info` → *how did control reach this logging call?* It renders the current call stack, headed `Stack (most recent call last):`, whether or not an exception exists. `stack_info` earns its keep on the puzzling non-exceptional log line: a warning that fires from a helper called in a dozen places, a deprecation notice, a "cache miss" you cannot attribute. It costs a stack walk plus formatting per call, so it belongs on rare lines, not on a hot path. ## Why not just format the traceback yourself `traceback.format_exc()` returns the same text as a string, and `log.error("failed: " + traceback.format_exc())` looks equivalent. It is worse in three ways. The traceback becomes part of the message string, so structured handlers cannot separate it; anything that filters, redacts or ships records field-by-field sees one opaque blob. Downstream formatters lose the ability to decide whether to render it at all. And, called outside a handler, `format_exc()` quietly yields the string `"NoneType: None"` and concatenates it into your message rather than being visibly absent. Let logging carry the exception as data and let the handler decide how to render it. ## Writing the message The traceback already tells you the exception type, its message and the code path. It does not know which of 6,800 rows you were on, which remote host you called, or which tenant the request belonged to. So the message should carry the *identifying* context and not restate the exception: `log.exception("inventory sync failed for batch %s at row %d", batch_id, offset)` beats `log.exception("got a TimeoutError")`. Use `%`-style placeholders with arguments rather than an f-string so the record keeps the parameters separately and the formatting cost is skipped when the level is disabled. ## The decision it implies Calling `log.exception()` is a claim: *this failure stops here, and this record is its report.* If you are going to re-raise, the exception is still travelling with its traceback intact and something above you will report it — logging here just doubles the report. That is the boundary rule this leaf is really about, and `logging.exception()` is the tool you use exactly once, at the layer that actually stops the failure.

  • What appears in the log if logging.exception() is called outside any except block?
    The message, followed by the line `NoneType: None`. `exc_info=True` captures the currently handled exception, and outside a handler there is none, so the record carries an empty `(None, None, None)` triple that the formatter renders that way. Seeing `NoneType: None` in production logs is a reliable signal that someone logged a traceback where no exception was live — often because the `except` block had already finished.
  • How would you attach a traceback to a WARNING rather than an ERROR?
    Pass `exc_info` to the level method directly: `log.warning("row deferred", exc_info=True)` inside the handler, or `log.warning("row deferred", exc_info=exc)` if you hold the exception object. `logging.exception()` is hardwired to ERROR, so using it just to get a traceback misrepresents severity — a failure you deliberately absorb and retry is a warning, not an error.
  • Why prefer log.exception('failed for %s', key) over an f-string in the message?
    With `%`-style placeholders the record keeps the message template and the arguments as separate fields, so aggregators can group thousands of records under one template, and the interpolation is skipped entirely when the level is disabled. An f-string is formatted eagerly at the call site and yields a unique string per record, which defeats grouping.

saying these in an interview costs you the question

  • Thinking logging.error() includes the traceback by default
  • Calling logging.exception() outside an except block
  • Believing exc_info only accepts True
  • Confusing stack_info with exc_info
  • Concatenating traceback.format_exc() into the message
  • Using logging.exception() purely to get a traceback at ERROR

context

open as a page

Why translate a third-party library's exception into your own exception class at a module boundary?

level: juniorimportance: must knowfreq 55%

basics

~20 s

So callers depend on your abstraction rather than on the library. If a database driver's error class escapes your function, every caller must import that driver to catch it, and replacing the driver breaks all of them.

open as a page

Why must an exception class raised in a child process be importable in the parent?

level: middleimportance: must knowfreq 45%

basics

~20 s

Pickle does not serialize a class body. It stores the class by reference -- module name plus qualified name -- and the receiving process re-imports that module and looks the name up. If the lookup fails there, your exception never reconstructs.

open as a page

Why does a published Python package define one base exception class that all its errors subclass?

level: middleimportance: must knowfreq 58%

basics

~10 s

So a caller can write one except PackageError to mean 'this library failed' without resorting to except Exception. The base marks the package's public error surface; the subclasses let callers who care be specific.

open as a page

Why does calling logger.exception() and then re-raising at every layer report one failure many times?

level: middleimportance: must knowfreq 50%

basics

~20 s

A re-raised exception keeps travelling with its full traceback, so every layer that logs it emits a complete report of the same failure. One error becomes N log records, N alerts and N counted errors. Log it once, where it stops.

open as a page

How would you write a Python retry decorator that retries only chosen exception types?

level: middleimportance: must knowfreq 65%

basics

~20 s

Loop over a fixed number of attempts, catch only a tuple of retryable exception classes, sleep on a growing delay between attempts, and re-raise on the last one. Decorate the wrapper with functools.wraps so it keeps the original function's name and docstring.

open as a page

When wrapping a library error, what does `raise MyError(...) from exc` preserve?

level: middleimportance: must knowfreq 50%

basics

~20 s

The original exception, attached to the new one as its __cause__. Both errors are then printed, so the domain type reaches the caller while the library's real message and stack frames stay available to whoever reads the log.

open as a page

When is a call safe for a retry decorator to repeat, and how do you make an unsafe one safe?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Only repeat calls whose second run leaves the same end state: reads, key-addressed overwrites, upserts. Appends, increments and sends are not. Make one safe by minting an operation key once, above the retry loop, so the far side can recognise the repeat.

open as a page

Can a parent process catch an exception raised inside a child process?

level: juniorimportance: should knowfreq 35%

basics

~20 s

No. A child is a separate OS process with its own memory and its own call stack, so the parent's try/except never wraps it. The exception object must be serialized and re-raised in the parent, or the parent learns only from the child's exit status.

open as a page

Why is calling sys.exit() from Python library code a design defect?

level: juniorimportance: should knowfreq 38%

basics

~20 s

sys.exit() raises SystemExit, which derives from BaseException, so a caller's except Exception never sees it and the whole process dies. Library code should raise its own exception and let the application decide whether to exit.

open as a page

Why does exponential backoff with time.sleep need random jitter and a delay ceiling?

level: middleimportance: should knowfreq 55%

basics

~20 s

Exponential growth gives a struggling dependency progressively more room instead of hammering it. Jitter randomises each wait so callers that failed at the same moment do not retry in lockstep, and a ceiling stops the doubling from turning into multi-minute sleeps.

open as a page

Why does a custom exception fail to unpickle with a TypeError about missing arguments?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Because BaseException.__reduce__ stores only the class and args, and unpickling reconstructs the object by calling the class with those args. If your __init__ signature does not match what you passed to super().__init__, that call fails.

open as a page

Why is the traceback missing when a child process's exception is re-raised in the parent?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Traceback objects reference live frames, code objects and locals, so they cannot be pickled. What crosses is the exception's class and args, and the parent's copy comes back with __traceback__, __cause__ and __context__ all set to None.

open as a page

When should a library call warnings.warn for a soft failure instead of raising an exception?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Warn only when the call still delivered what it promised and the caller should change something next time. An incomplete or wrong result — a silently truncated document — is a failure: raise it, or state it in the return value.

open as a page

An inventory sync hits intermittent timeouts partway through a 6,800-row batch — what do you log, re-raise or swallow?

level: seniorimportance: should knowfreq 45%

basics

~10 s

Decide per scope: a single row's timeout is absorbed, counted and logged once at WARNING; the batch raises a failure once deferred rows cross an agreed threshold. Never exit successfully after swallowing everything.

open as a page

Should a Python retry loop wrap the entire `with` block or live inside it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It depends on whether the failure damaged what the block is holding. If the resource carries protocol or transaction state, wrap the whole with statement so each attempt acquires a fresh one; retrying inside reuses something the failure may already have poisoned.

open as a page

Which layer should translate a database driver's error in a nightly report generator with a storage adapter, a service layer and a CLI?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The storage adapter - the only module that imports the driver. It maps driver failures onto domain error types; the service layer handles those types without knowing the driver exists, and the CLI turns whatever escapes into a message and an exit status.

open as a page

How would you evolve a published library's exception hierarchy without breaking existing except clauses?

level: principalimportance: should knowfreq 30%

basics

~20 s

Only add downward. A new class that subclasses an already-documented type keeps every existing handler working; removing a class, dropping a base, or raising a different type from the same function silently unhandles code you cannot see.

open as a page

How do you set a codebase-wide policy for which layers log exceptions and which re-raise?

level: principalimportance: should knowfreq 34%

basics

~10 s

Name a small, explicit set of boundary layers - request entry, worker loop, thread or process entry, CLI main - as the only places allowed to log an exception and stop. Everything else raises.

open as a page

Where should retry logic live in a Python codebase, and when is a hand-rolled decorator the wrong choice?

level: principalimportance: should knowfreq 40%

basics

~20 s

Put one retry layer on each call path, in the thinnest adapter that still understands the transport's exceptions. Nested decorators multiply attempts. Hand-rolled is fine for one policy; a maintained library earns its place once you need async variants, deadline propagation and telemetry.

open as a page

When is translating every library exception into a domain hierarchy the wrong call for a 4-person team?

level: principalimportance: should knowfreq 33%

basics

~20 s

When the mapping costs more than the coupling it removes: one application, one team, one library that will not be replaced. Translation earns its keep at boundaries someone else calls or where the dependency is genuinely swappable - not at every internal module edge.

open as a page