What does "predictable error semantics" mean under the Principle of Least Astonishment, and what are the most common ways an API surprises callers when things go wrong?
answer
- one signalling mechanism per failure category
- classify: caller error / absence / conflict / transient / bug
- state after failure: strong guarantee = unchanged
- retryable? idempotency key for writes
- never swallow a failure into a default value
basics
~20 sIt means failures are reported the same way everywhere, so callers can predict what happens without reading each implementation: one style of signalling errors, consistent handling of "not found" versus "invalid" versus "broken", and a clearly defined state after a failure.
solid answer
~50 sPredictable error semantics means the failure contract is uniform and explicit. Four things must be consistent. **Signalling mechanism**: don't mix exceptions, null returns, sentinel values, and result types for the same class of failure in one API. **Classification**: distinguish caller mistakes (invalid input, unauthorized) from expected absence (not found) from infrastructure faults (timeout, unavailable), and map each to one representation. **State after failure**: say whether the operation is atomic (nothing happened), partially applied, or leaves the object unusable — ideally guarantee that failure means unchanged. **Retryability and idempotency**: callers must be able to tell from the error whether retrying is safe and useful. Beyond that: don't leak internal error types across a module boundary, don't swallow errors into default values, don't reuse one generic error for everything, keep HTTP/RPC status mapping stable, and never disclose secrets in messages. Surprising failures are worse than surprising successes because they surface under pressure.
code
pseudo · 14 lines// Astonishing: three mechanisms, no classification, unclear state
findUser(id) -> User | null // absence via null
getOrder(id) -> throws NotFound // absence via exception
loadInvoice(id) -> Invoice | EMPTY // absence via sentinel
transfer(a,b,amt) -> throws Exception // did the money move? unknown
// Predictable: one rule per category, documented state
findUser(id) -> Optional<User> // absence is normal -> empty
requireUser(id) -> User | NotFoundError // absence is a caller error here
transfer(a, b, amount, idempotencyKey)
-> Ok | InvalidAmount (permanent, do not retry)
| InsufficientFunds (permanent, do not retry)
| Unavailable(retryAfter) (transient, safe to retry with same key)
// guarantee: any failure above leaves both balances unchangedgo deeper
Say errors should be reported the same way everywhere, and that mixing null returns with exceptions for the same situation confuses callers.
Add classification (invalid input vs not found vs transient) and note that swallowing errors into defaults hides bugs.
Cover the full failure contract — mechanism, classification, state after failure, retryability and idempotency, boundary translation — plus the trade-off between result types and exceptions.
Treat the error taxonomy as a cross-team contract: stable codes, versioning and deprecation of statuses, retry/idempotency policy interacting with client behavior and incident blast radius, and enforcement through platform libraries and contract tests.
## Terms - **Signalling mechanism**: how an operation reports failure — a thrown exception, a returned `null`/optional, a sentinel value like `-1`, an error code, or a result/either type carrying success-or-error. - **Checked vs unchecked** (where a language distinguishes them): whether the compiler forces the caller to acknowledge a failure. - **Exception safety levels** (from C++ practice, applicable anywhere): *no-throw* (never fails), *strong* (failure leaves state exactly as before — commit-or-rollback), *basic* (state stays valid but may have changed), *none* (state may be corrupt). - **Idempotent**: repeating the call yields the same observable result — the precondition for safe retries. - **Fail-fast**: detect and report invalid state at the earliest point rather than continuing and failing later somewhere unrelated. ## The five surprises **1. Inconsistent signalling.** In one API, `findUser` returns null, `getOrder` throws, `loadInvoice` returns an empty list, and `fetchAccount` returns a status code. Callers cannot form a habit, so every call site is written by trial and error and half handle the wrong case. Pick one mechanism per failure category and hold the line. **2. Missing classification.** Collapsing everything into one generic error (or one HTTP 500, or one `ServiceException`) forces the caller to string-match messages to decide what to do. At minimum separate: *caller error* (invalid input, precondition violated, unauthorized), *absence* (not found — often not an error at all), *conflict* (concurrent modification, duplicate), *transient* (timeout, unavailable, throttled), and *bug* (invariant violated). The remedy is a small closed set of error codes or types with a documented mapping to transport-level statuses. **3. Undefined state after failure.** If `transfer()` fails, was money moved? An API that cannot answer this is unusable. Prefer the **strong guarantee**: validate first, mutate last, so failure means nothing changed. Where partial application is unavoidable (a batch operation), report *which* items succeeded rather than a bare failure, and make the operation resumable. **4. Unclear retryability.** Callers with a retry policy must distinguish "retry will never help" (invalid input) from "retry after backoff may help" (timeout, throttle, unavailable). Ambiguity produces both extremes: retry storms that amplify an outage, and no retry where one would have healed. Mark transient failures explicitly, expose backoff hints where you have them, and pair retryable writes with an idempotency key so a retry cannot double-charge. **5. Leaky or lying errors.** Internal exception types, driver-specific errors, or stack traces crossing a module or service boundary make consumers depend on your internals and break them when you swap implementations. The mirror-image sin is **swallowing**: catching broadly and returning a default (empty list, zero, `false`), which converts a loud failure into silent wrong data — the most astonishing outcome of all. Also avoid exceptions as ordinary control flow, and never put secrets or personal data into messages that reach clients or logs. ## Making it predictable in practice - Write the failure contract next to the signature: what can fail, how it is signalled, what state results, whether retry is safe. - Keep transport mappings stable — if "not found" is 404 today it must not become 200-with-empty-body tomorrow; that breaks every client. - Translate at boundaries: catch infrastructure errors in the adapter layer and re-express them in the module's own vocabulary, preserving the cause for diagnostics. - Use the type system where you can — an optional return says "absence is normal", a result type says "failure is expected and must be handled", a thrown exception says "this is exceptional". - Make messages actionable: what failed, which input, what to do next, plus a correlation id for support. - Test the failure paths. Untested error handling is where astonishment breeds, because nobody has ever observed it before the incident. ## Trade-offs Explicit result types make failure impossible to ignore but add noise and push handling into every layer. Exceptions keep the happy path clean but hide the failure contract from the signature in most languages. Neither is universally right; **consistency within one system matters more than which one you pick**, because the whole point is that a caller who has learned one call site has learned them all.
- Should "entity not found" be an exception or an empty optional?It depends on whether absence is expected at that call site. For a lookup where missing is a normal outcome (checking an email during registration), an empty optional is honest and cheap. For a call where the id was already proven to exist, absence means a broken invariant and an exception is right. The astonishing choice is doing both inconsistently across one API.
- How do you keep error semantics predictable across a service boundary, where the caller is a different team?Publish a stable, versioned error contract: a closed set of machine-readable codes, a documented mapping to transport statuses, an explicit retryable flag or retry-after hint, and a correlation id. Never expose internal exception class names or stack traces. Treat a code or status change as a breaking change and cover failure paths in contract tests.
An airline that announces delays in minutes at one gate, in local time at another, and not at all at a third. Every passenger relearns the system at every gate, and the ones who guess wrong miss the flight.
saying these in an interview costs you the question
- Catching broadly and returning an empty collection or zero so the call "never fails"
- Using one generic error type for everything and forcing callers to parse messages
- Letting persistence or driver exceptions escape a module's public API
- Retrying every failure indiscriminately, including permanent input errors
- Returning HTTP 200 with an error object for genuine failures, or 500 for a validation problem
- Assuming a failed write left nothing changed without ever documenting or testing it
- Putting internal details or secrets into messages returned to clients