skip to content

How granular should a module's custom exception subclasses be?

level: seniorimportance: should knowfreq 48%

answer

  1. Count handlers, not failure modes
  2. A class per decision, a field per detail
  3. Write the except clause first
  4. Intermediate bases express retryability
  5. Coarse types invite catch-and-continue

basics

~20 s

Split when a caller would handle the cases differently; otherwise keep one class and put the detail in an attribute. A class per decision, a field per detail: a distinction no handler branches on is noise.

solid answer

~50 s

The unit of granularity is the **caller's decision**, not the number of ways your code can fail. A new subclass earns its place when catching it would change behaviour — retry, skip this line, fail the run — and otherwise the distinction belongs in an attribute. Keep the module base class as the broad handle and never raise it directly. When the decision axis is not "what failed" but "is retrying worth it", express that as intermediate base classes — a transient branch and a permanent branch — so a caller can catch the branch instead of enumerating leaf types that you will keep adding. Both extremes hurt: a single catch-all pushes callers into `except PickListError: pass`, which silently swallows failures that needed cleanup, while a class per raise site freezes your internal structure into your public API and gets caught at the base anyway.

code

python · 26 lines
python
class PickListError(Exception):
    """Anything this module raises."""


class TransientPickError(PickListError):
    """Retrying may succeed."""


class PermanentPickError(PickListError):
    """Retrying will not help."""


class InventoryTimeoutError(TransientPickError):
    pass


class SkuNotFoundError(PermanentPickError):
    pass


def react(exc: PickListError) -> str:
    return "retry" if isinstance(exc, TransientPickError) else "skip line"


print(react(InventoryTimeoutError("north")))
print(react(SkuNotFoundError("A-1")))

go deeper

for a junior

Know that catching a base class catches all its subclasses, and that the hierarchy a library exposes is what decides how precisely you can handle its failures.

for a middle

Explain the tradeoff in both directions: what a single catch-all type does to callers, and why a class per raise site collapses back into everyone catching the base anyway.

for a senior

Show the diagnosis: recognising that catch-and-continue handlers with skipped cleanup trace back to a taxonomy with no branch structure, and designing transient versus permanent bases so retries are expressible.

for a principal

Own the hierarchy as a compatibility surface — where new types may be added safely, what a rename costs across consumers, and when a long tail moves from classes to codes so the tree stays small.

### Count the handlers, not the failures Every exception class you export is a question you are asking callers: *do you want to treat this case specially?* If no realistic caller would answer yes, the class is overhead — one more name in the API, one more thing to document, one more import. So the design rule is short: **a class per decision, a field per detail.** Before adding a subclass, write the handler that would catch it. If that handler does something different from the handler for its sibling, the subclass is real. If it does the same thing and only logs a different word, the difference is data. ### A worked example: the warehouse pick-list builder A pick-list builder turns an order into a sequence of aisle locations, calling an inventory service per line. It can fail in a dozen ways, but callers only make three decisions: **skip this line and carry on**, **retry the call**, or **abandon the run**. That is the shape of the hierarchy, whatever the internal failure catalogue looks like. A first version shipped with exactly one type, `PickListError`, and every failure raised it with a different message. Callers did the only thing a single type allows: `except PickListError: continue`. That handler was correct for a missing SKU and wrong for everything else, and because a timeout also unwound the block that released the run's reservation handle, every swallowed timeout left a reservation unclosed. Nothing showed up in normal operation. It surfaced after a deploy, when a 45-second cold start on the inventory dependency made timeouts common for the first minute of every restart and the handle pool ran dry. The bug was not the missing cleanup alone; it was a taxonomy that gave callers no way to tell "skip it" from "this call never happened". The fix was two intermediate classes, not twenty leaf ones: ```python class PickListError(Exception): ... class TransientPickError(PickListError): ... class PermanentPickError(PickListError): ... ``` Every concrete error derives from one branch. A caller writes `except TransientPickError: retry()` and `except PermanentPickError: skip_line()` and is still correct next quarter when three new leaf types appear, because new leaves join an existing branch. ### The other failure mode: too many classes Over-slicing is quieter but just as real. A module with one exception class per `raise` statement — twenty-three of them, named after the internal function that raised — produces an API nobody can hold in their head, so callers catch the base class and branch on nothing. Worse, those names encode your implementation: rename the internal step and either the class name lies or you break every handler. Symptoms: classes that differ only in message wording, classes that appear in exactly one `raise` and no `except` anywhere in the organisation, and documentation that lists error types alphabetically because there is no structure to present. ### Where the long tail goes Real systems have a long tail of variants that nobody branches on but everybody wants to see: a reason code, a field name, an upstream status. Put those on the instance — an attribute, ideally an enum member — rather than minting a class each. Codes scale to hundreds of values, are easy to aggregate in logs and dashboards, and can be added without touching the class tree. Their weakness is that they do not compose with `except`, which is exactly why the handful of decisions stay as classes. ### Two more levers worth knowing **Inherit a built-in where the meaning genuinely matches** — a lookup failure that is also a `KeyError`, a bad argument that is also a `ValueError` — so callers who never read your docs still catch the right thing. Do not do it to broaden the net; you inherit the built-in's promise too. **Keep the base abstract by convention.** Raising the base directly is a smell: it means a failure mode was never classified, and callers cannot discriminate. If you find yourself doing it, that is the signal that a new decision — and therefore a new subclass — exists. ### How to talk about it in an interview Start from the handler, not from the failure list. Say out loud that the hierarchy is public API: adding a leaf under an existing branch is safe, adding a top-level sibling is a change every broad handler misses, and renaming anything is a break. That framing — decisions, branches, and the compatibility cost of moving a class in the tree — is what separates a senior answer from a list of error names.

  • When is an error-code attribute a better choice than another exception subclass?
    When the variants are a long tail nobody branches on — upstream status values, validation reason codes, field names. Codes scale to hundreds of values, aggregate well in logs, and can be added without touching the class tree. The tradeoff is that a code cannot be matched by an `except` clause, so anything a caller actually handles stays a class.
  • How do you add a new failure mode without breaking existing handlers?
    Derive it from an existing subclass or branch, so every handler that already catches that branch keeps matching. Introducing a new top-level sibling under the base is the risky move: only handlers catching the base see it, and anyone catching specific branches now misses a case. That asymmetry is why the branch layer is worth having up front.
  • Is it ever right for a module to raise its own base exception class directly?
    Effectively no. The base is a catch handle, and raising it tells a caller only "something in here failed", which forces the coarse handler you were trying to avoid. When you are tempted, it usually means a failure mode has not been classified yet — which is the signal that a new subclass is warranted.

Exception classes are the buttons on a control panel: one per action an operator can take, not one per thing that can go wrong inside the machine.

saying these in an interview costs you the question

  • One catch-all error type for an entire package
  • A distinct exception class for every raise statement
  • Encoding the case in the message for callers to parse
  • Catching the module base class and continuing regardless
  • Naming exception classes after internal implementation steps

context