skip to content

When should a Python API raise a custom exception rather than return a result value?

level: principalimportance: should knowfreq 38%

answer

  1. Who is expected to see this failure?
  2. Exceptions are loud, ignored returns are silent
  3. Batch work reports failures as data
  4. One aggregate beats five thousand raises
  5. The choice is API surface, not style

basics

~10 s

Raise when the failure is outside the caller's expected outcomes and most callers cannot continue; return a value when failure is a routine result every caller inspects anyway. Batch work aggregates failures instead.

solid answer

~50 s

The standard library models both answers: indexing a dict raises `KeyError`, `dict.get` returns `None`, and the difference is who expects the miss. Raise for contract violations and conditions the caller did not ask about — a raise propagates by default, so nobody accidentally continues on bad state. Return a value when absence or failure is an ordinary outcome and forcing a `try` around every call would drown the logic; the cost is that a returned failure is silently ignorable and usually explodes later, far from its cause, unless the type system makes it visible. For batch APIs, per-item failures are data: collect them and let the caller see all of them, optionally raising an `ExceptionGroup` (3.11) for the aggregate rather than aborting on the first bad row. Whichever you choose is API surface — flipping a raise into a return breaks callers exactly as a signature change would.

code

python · 25 lines
python
class SkuNotFoundError(Exception):
    def __init__(self, sku):
        super().__init__(sku)
        self.sku = sku


CATALOGUE = {"A-1": "aisle 3", "B-2": "aisle 7"}


def build(skus):
    picked, failures = [], []
    for sku in skus:
        try:
            picked.append(CATALOGUE[sku])
        except KeyError:
            failures.append(SkuNotFoundError(sku))
    if failures:
        raise ExceptionGroup("pick list incomplete", failures)
    return picked


try:
    build(["A-1", "Z-9", "B-2", "Q-0"])
except* SkuNotFoundError as group:
    print([e.sku for e in group.exceptions])

go deeper

for a junior

Recognise the two styles in code you read: an operation that raises on failure versus one that returns None or a default, and know that catching or checking is not optional either way.

for a middle

Explain why an unhandled exception fails loudly while an ignored return fails far from its cause, and give the standard-library pair that shows both designs on the same operation.

for a senior

Argue the batch case concretely: per-item failures as data, an aggregate at the end, and what a per-row raise does to logs and rerun cost when a dependency is briefly unavailable.

for a principal

Own the rule rather than the case — which layers raise and which return, where failures are aggregated and converted, and the fact that flipping a raise into a return breaks callers silently and therefore counts as a breaking change.

### The real question: who is expected to see this? Both designs are legitimate in Python, and the standard library uses both deliberately. `d[k]` raises `KeyError`; `d.get(k)` returns `None`. Same operation, two expectations: indexing says "this key is part of my contract", `get` says "absence is a normal answer". So the decision is not about style, it is about which outcomes are inside the caller's plan. **Raise when** the condition means the caller's assumptions were wrong, when most callers cannot sensibly continue, or when the failure is rare enough that forcing every call site to check would be pure noise. **Return a value when** failure is an expected outcome of the operation, when the caller almost always has something to do with it, or when the call sits in a hot loop where most iterations fail and unwinding per iteration is genuinely measurable. ### The asymmetry that decides most cases An unhandled exception is loud: the process stops, the traceback names the line, and the failure is impossible to ignore accidentally. An ignored return value is silent: the `None` flows onward and surfaces two hundred lines later as an `AttributeError` on something that was never the problem. That asymmetry is why exceptions are the default for anything a caller might forget, and why returning failure only works when the type makes the failure impossible to overlook — an `Optional` return that a type checker enforces at the call site, or a result object whose value cannot be read without acknowledging the failure branch. The mirror image is also a real defect: raising for outcomes every caller expects. An API that raises on a cache miss forces `try` around code that is not exceptional at all, and eventually somebody writes a bare `except` there and swallows the real errors with it. ### The batch case, where the answer is usually "both" A pick-list builder processing 5,000 order lines against an inventory service is the archetype. If a single unknown SKU raises, the whole run dies and the 4,999 good lines are lost; worse, the operator learns about one bad row per attempt and has to rerun to find the next. The useful design reports per-line failures as data — a result carrying picked locations and a list of failure objects — so one pass shows every problem. An aggregate exception is the right tool when the *call as a whole* failed and you still want all the reasons. Since 3.11, `ExceptionGroup` (with `except*` for handling) exists exactly for this: one raise carrying many independent failures, each a real exception with its own payload. Before 3.11 the same job needed a custom exception holding a list, which is still fine, and is still the choice when the group semantics buy nothing. Operational reality sharpens this. During a 45-second cold start the inventory dependency answers nothing, and a builder that raises per line produces 5,000 stack traces, five thousand log lines and no summary; a builder that aggregates produces one error saying "5,000 of 5,000 lines failed, all transient" — which is both cheaper and actionable. ### Constraints that push the decision * **Distance between the raiser and the handler.** Exceptions cross layers for free, which is a virtue when the caller three frames up is the only one who can decide, and a liability when the failure is meaningful only locally. If nobody outside can act on it, handle it where it happens and return an ordinary outcome. * **Reconstructibility.** An error crossing a worker boundary has to be rebuilt on the other side, so a design that raises heavily depends on exception payloads that survive the trip. A design that returns plain data avoids that class of problem entirely. * **Cost on a failing hot path.** Setting up a `try` is effectively free on the happy path since 3.11, but actually raising is not. If a loop's common case is failure, a sentinel return can be the honest choice — measure before rewriting. * **Consistency with the ecosystem.** A house style where every function returns a `(value, error)` pair fights the standard library and every library you compose with; you end up converting at every boundary. Pick the idiom Python callers expect, and reserve returned-failure APIs for the places that clearly earn it. ### Why this is a lead's call The error contract is API surface with the same weight as a signature. Turning a raise into a returned `None` breaks every caller that relied on propagation, and it breaks them *silently* — no import error, no type error at runtime, just data flowing past a check nobody wrote. So the decision belongs in the design review, is documented next to the function, and stays stable across versions. What a lead owns is the rule, not the individual case: which layers raise, which return, where aggregation happens, what every error type carries, and how failures are converted once at each boundary rather than ad hoc in twenty call sites.

  • How do you stop a returned failure value from being silently ignored?
    Make it visible in the type — an `Optional` return or a result object a type checker forces the caller to unpack — and make the success value unreachable without touching the failure branch. Documentation and review help, but the durable mechanisms are static checking and a shape that cannot be used incorrectly by accident. If neither is available, raising is the safer default.
  • Where does ExceptionGroup fit in the raise-versus-return decision?
    It is the middle option for a call that genuinely failed but has several independent reasons. Since 3.11 you can raise one group carrying every failure, each with its own payload, and callers handle subsets with `except*`. It keeps the loudness of an exception while avoiding first-failure-wins, which is what makes batch runs painful to debug.
  • Does a hot loop change the answer?
    Sometimes. Entering a `try` costs nothing measurable since 3.11, but actually raising and unwinding does, so a loop whose common case is failure can be meaningfully faster with a sentinel-returning API. Treat it as a measured optimisation for a specific hot path, not as a general argument against exceptions.

Raising is pulling the fire alarm — everyone downstream hears it whether they wanted to or not. Returning a failure value is leaving a note on the desk: cheaper, and only useful if someone reads it.

saying these in an interview costs you the question

  • Returning None for every failure so nothing ever raises
  • Raising for outcomes every caller expects, such as a cache miss
  • Claiming exceptions are always too slow for real code
  • Aborting a five-thousand-row batch on the first bad row
  • Treating the error contract as an internal implementation detail

context