How would you design a consistent error-handling strategy across a multi-service system so failures are diagnosable, retryable where safe, and never silent?
answer
- Taxonomy that drives action: input / business / transient / bug
- Stable codes in the contract; never branch on messages
- One boundary mapper: status + code + safe body + one log line
- Retryable only with idempotency; backoff+jitter, one retry layer
- Async: DLQ + idempotent consumers; correlation id; log once
basics
~20 sAgree one error taxonomy (client error / business rejection / transient infrastructure / bug), give each a stable code and a retryable flag, translate to transport in one place per service, log once with a correlation id, and make retries safe by requiring idempotency.
solid answer
~50 sStart from a shared **taxonomy** because it drives behaviour: invalid input (never retry, 4xx), business rejection (never retry, needs a human/user decision), transient infrastructure failure (retry with backoff and jitter, 5xx/503), and programmer error (fail the unit of work, page someone). Publish it as **stable error codes** in the API contract — clients branch on codes, never on message strings, and codes are versioned like any contract element. Each service maps domain errors to transport in **one boundary handler** (problem+json / gRPC status / house envelope), producing a safe client message and a full internal log entry with a correlation/trace id — logged once, at the boundary, with the cause chain intact. Retries require idempotency keys, otherwise 'safe to retry' duplicates side effects; add timeouts, circuit breakers and bulkheads so a slow dependency doesn't cascade. Finally, make silence impossible: alert on error-rate SLOs, treat empty catch blocks as a lint failure, and inject failures in tests.
code
pseudocode · 20 lines// one boundary mapper per service
onError(err, request) {
e = classify(err); // INPUT | BUSINESS | TRANSIENT | BUG
logOnce(level(e), {
code: e.code, traceId: request.traceId,
dependency: e.dependency, cause: chainOf(err) // full chain, internal only
});
respond({
status: statusFor(e), // 400 / 409 / 503 / 500
code: e.code, // stable, documented, client branches on this
message: safeMessage(e), // no stack, no SQL, no PII
retryable: e.retryable,
traceId: request.traceId
});
}
// retry policy only where side effects are idempotent
retry(attempts=3, backoff=exponentialWithJitter, budget=callerDeadline) {
charge(order, idempotencyKey = order.id)
}go deeper
Keep it concrete: consistent error codes, don't swallow errors, log with an id you can search, and return a helpful but non-revealing message to the user.
Add the taxonomy and its actions, one boundary handler per service, wrapping with preserved causes, and retry with backoff for transient failures only.
Bring in idempotency for safe retries, backoff+jitter and single-layer retry, circuit breakers/bulkheads/deadline propagation, async DLQ handling, and log-once with correlation ids.
Frame the error model as a versioned public contract with stable codes and retryability semantics; discuss organisational rollout (standardise semantics, not a shared library), degradation policy per dependency, SLO-based alerting, and enforcement via lint, fault injection and runbooks.
## Why a *strategy* rather than local good manners In one class, clean error handling is a readability concern. Across a system it becomes an operational contract: whether a failure is retried, whether it wakes someone at 3am, whether the user sees something actionable, and whether an engineer can reconstruct what happened. Left to individual judgement, every service invents its own conventions and the composite behaviour becomes unpredictable — one service retries a business rejection forever, another swallows a timeout and returns a 200 with an empty body. ## 1. A shared taxonomy that drives behaviour The categories are only useful if each implies a different action: | Category | Cause | Retry? | Transport | Who is woken | |---|---|---|---|---| | Invalid input | Malformed/failing validation | No — deterministic | 400 / INVALID_ARGUMENT | Nobody | | Business rejection | Insufficient funds, out of stock, forbidden | No — needs a decision | 409/422/403 | Nobody | | Transient infrastructure | Timeout, connection reset, 503, deadlock | Yes, bounded, with backoff | 503 / UNAVAILABLE | On rate/SLO breach | | Programmer error / invariant broken | Null deref, impossible state | No — fix it | 500 / INTERNAL | Yes | | Dependency degraded | Downstream over budget | Maybe: degrade or shed | 503 + Retry-After | On SLO breach | The two attributes worth encoding explicitly on every error are **retryable** and **client-visible**. Everything else is diagnostics. ## 2. Stable error codes as part of the contract Clients must be able to branch programmatically. That means a stable, documented code (`PAYMENT_DECLINED`, `RATE_LIMITED`) — not an HTTP status alone (too coarse), not a message string (localised, reworded, breaks silently), not an internal exception class name (leaks structure and changes on refactor). Practical rules: codes are additive-only within a major version; a client seeing an unknown code must have defined default behaviour (usually: treat as terminal, surface generic message); document each code with cause, retryability and recommended client action. Reserve a machine-readable payload for details (which field, which limit, when to retry). ## 3. One translation point per boundary Each service should have exactly one place that turns internal failures into an outward representation: an exception-mapping handler / middleware / interceptor. It: - maps exception or Result-error type → status + code + safe message; - attaches the correlation id so the client can quote it in a support ticket; - emits **one** structured log entry at the appropriate level, with the full cause chain; - strips anything sensitive — stack traces, SQL, internal hostnames, PII, secrets. Verbose error output is a real reconnaissance vector, and stack traces to clients are a standard finding in security reviews. Symmetrically, at the *inbound* edge of each adapter, third-party and transport errors are wrapped into domain terms so the core never sees `SQLException` or an SDK type. ## 4. Retries only where they're safe "Retryable" is a claim about *side effects*, not about the error alone. A timeout is the hard case: the request may have succeeded and you lost the response. Retrying a non-idempotent write duplicates the charge. So the strategy must pair retryability with **idempotency keys** (client-supplied id the server deduplicates on) or naturally idempotent operations (`PUT`, conditional updates with expected-version). Then: - exponential backoff **with jitter**, or synchronised clients retry in lockstep and create a thundering herd; - a bounded attempt count and a total time budget derived from the caller's deadline (retrying past the caller's timeout is pure waste and amplifies load); - **retry at one layer only** — retries at client, gateway, and service multiply (3×3×3 = 27 downstream calls) and are a classic cause of retry-storm outages; - a **circuit breaker** so a sustained failure stops generating traffic and fails fast instead; - **bulkheads** (separate pools/queues per dependency) so one slow dependency can't consume all threads; - **load shedding** and deadline propagation so an overloaded service rejects early instead of timing everything out. ## 5. Asynchronous paths need their own model Queues and event consumers change the shape: failures are not returned to a caller. You need a redelivery policy, a maximum attempt count, a **dead-letter queue** with enough context to replay, poison-message handling, and — because at-least-once delivery is the norm — idempotent consumers. Silent dead-lettering is the async equivalent of an empty catch block, so DLQ depth must be alarmed and owned. ## 6. Diagnosability - **Correlation/trace id** generated at the edge, propagated on every hop and every log line, so one identifier reconstructs the whole request path. - **Log once.** Log-and-rethrow at each layer multiplies noise and hides the real handler. Wrap with context as it rises; log where it is finally handled. - **Structured logs** (fields, not sentences) so error rates can be aggregated by code and dependency. - **Errors as metrics**: rate by code, by dependency, by client — alerting on SLO burn rather than on individual exceptions. - **Preserve the cause chain** end to end; a code with no provenance is only half the story. ## 7. Failure of the *system*, not the request Decide deliberately, per dependency, between fail-fast and degrade: is a recommendation service optional (render the page without it) or essential (fail the request)? Write it down — the default of "whatever the code happens to do" produces user-visible outages from optional features. Graceful degradation, sensible defaults, cached fallbacks, and read-only modes are design decisions, not accidents. ## 8. Making silence impossible - Lint/static analysis rules that fail the build on empty catch blocks, swallowed errors, ignored Results, and bare broad catches. - Tests that exercise failure paths; fault injection or chaos experiments for timeouts, 500s, and slow dependencies. - Runbooks keyed by error code so an alert has an action attached. - Review checklist item: for every new failure mode, what is the code, is it retryable, who is paged, what does the user see? ## Trade-offs to state out loud A rigid global taxonomy can be over-engineering for a small system, and forcing every team through one shared error library creates a coupling point and a release bottleneck. The stable compromise: standardise the **contract** (codes, retryability, transport mapping, correlation id, log-once) and let each service implement it with its own idioms. Standardise semantics, not code.
- A client times out on a payment request. It doesn't know whether the charge went through. How does your strategy handle it?By requiring an idempotency key on the write: the client retries with the same key, and the server either returns the original result (if it committed) or performs the charge exactly once. Without that key, 'retryable' is a lie for any non-idempotent operation — timeouts are ambiguous by nature, so the safety must come from the operation's design, not from the error classification.
- Why is retrying at multiple layers dangerous?Retries multiply. Three attempts at the client, gateway and service turn one user request into up to 27 downstream calls, so a partially degraded dependency gets hit with an order of magnitude more load exactly when it is weakest — a retry storm. Pick one layer to own retries, cap total attempts against the caller's deadline, add jitter, and put a circuit breaker in front.
- How do you keep this from becoming bureaucratic over-engineering in a small system?Standardise semantics, not implementation: a short list of categories, a retryable flag, stable codes, one boundary mapper, a correlation id, and log-once. That is a page of policy, not a shared framework. Introduce dead-letter queues, breakers and bulkheads when there are asynchronous paths and real dependency fan-out to justify them.
- What belongs in the client-facing error body versus the internal log?Client: stable code, human-safe message, correlation id, retryability, and structured details it can act on (which field, retry-after). Internal: the full cause chain, stack, dependency, query/request identifiers, and timing. Never leak stack traces, SQL, internal hostnames, secrets or PII outward — it's both a support burden and a reconnaissance aid.
Air-traffic control: every incident gets a standard classification and a defined response, one identifier follows the flight through every handoff, and the recording keeps the full chain. Individual controllers improvising their own categories is precisely how mid-airs happen.
saying these in an interview costs you the question
- Clients branching on error message strings or on exception class names
- Marking an operation retryable without an idempotency mechanism
- Retrying at every layer, producing multiplicative load during an outage
- Returning stack traces or SQL to clients
- Log-and-rethrow at every layer, so one failure appears five times with no single authoritative entry
- Dead-letter queues with no owner, no alarm and no replay path
- Using HTTP status alone as the error contract — too coarse for clients to branch on
- A shared company-wide error library treated as mandatory, becoming a coupling and release bottleneck