skip to content

How would you decide which hooks in a shared middleware chain may recover from errors rather than let them unwind?

level: principalimportance: nice to knowfreq 33%

answer

  1. one translation point
  2. observe and rethrow by default
  3. optional concerns may degrade
  4. mandatory checks never recover
  5. every recovery must be counted

basics

~20 s

Default to one place that turns failures into responses and let every other hook observe and rethrow. Permit recovery only where the concern is optional, the degraded outcome is defined, and the recovery is counted rather than silent.

solid answer

~40 s

Start from a single translation point — one outer stage that decides what a failure means to the caller — because two of them produce responses that depend on which path a failure took. Everything between that stage and the handler may **observe** a failure, tagging or logging it, and must rethrow. Recovery is the exception, and it needs three conditions: the concern is genuinely optional, the degraded result is defined and still correct, and the recovery is observable through a counter or log so a persistent inner failure is not hidden. Mandatory concerns — identity, permission, tenancy, input validation — may never recover, because continuing after they fail converts a refusal into a silent grant. Cleanup that releases a resource is not recovery and must still let the failure continue.

go deeper

for a junior

The safe default is to let a failure travel outward. If you catch one in a hook, you have quietly decided the request succeeded, which is rarely yours to decide.

for a middle

Distinguish three roles clearly: passing through, observing and rethrowing, and recovering. Only the third changes the outcome, and it needs a stated reason.

for a senior

Argue the conditions for recovery on a concrete hook and show how the degradation is measured, so a permanently failing dependency is noticed rather than absorbed.

for a principal

Own the estate-wide rule and its enforcement: one translation point, a chain test that proves it, counters on sanctioned degradation, and alerts on those counters.

Once a chain is shared across many teams, "who may catch?" stops being a coding-style question and becomes an architectural policy. Every catch that produces a result instead of rethrowing is a decision that the request succeeded, taken by an element that usually knows the least about what success means. The job of the policy is to keep those decisions few, deliberate and visible. ## Start from one translation point The default is that exactly one element — the outer error stage — converts failures into responses. The reason is not purity but debuggability: with two translation points, the answer a caller receives depends on how deep in the chain the failure happened, and the same underlying fault presents two different ways. One translation point also concentrates the decisions that are easy to get wrong, so they can be reviewed in one file rather than discovered across dozens of hooks. Everything inside that stage falls into one of three roles: - **Pass-through.** It does its work and does not interact with the failure path at all. - **Observer.** It catches, records or enriches, and rethrows. The outcome is unchanged. - **Recoverer.** It catches and produces a result. The failure stops there, and this is the role that needs a licence. ## The questions that decide whether a hook may recover 1. **Is the concern optional?** If the request is still correct without this element's contribution, degradation is legitimate. If the element exists to say no, it is not. 2. **Is the degraded outcome defined?** "Continue with whatever we have" is not a definition. "Return the response without the enrichment block" is. 3. **Is the recovery visible?** A counter or a log line naming the recovering element. Recovery that produces no signal cannot be distinguished from the system working. 4. **Who owns the resulting response?** If the answer is the recovering hook, it has taken over a job the translation point owns, and that needs to be explicit rather than incidental. 5. **Would a permanent failure here be noticed within an hour?** If not, the recovery is a future outage that nobody will connect to its cause. ## A classification that scales across teams | Hook class | On failure | Why | |---|---|---| | Identity, permission, tenancy resolution | never recover | continuing turns a refusal into a silent grant | | Input parsing and validation | never recover | the handler would receive input nobody vouched for | | Optional enrichment, caching, personalisation | may recover | the response is plainer, not wrong | | Observability: logging, metrics, tracing | observe and rethrow | it must never change the outcome it measures | | Resource acquisition and release | cleanup only | releasing is not deciding what the caller is told | The security rows are the ones worth being absolute about. A permission check that fails and is then recovered from produces a request that proceeds without the check ever having passed, and it produces it silently — the worst combination available. ## Keeping the policy alive A policy that lives only in a document erodes at the pace of new hires. The cheap mechanisms are: - An **assembled-chain test** that runs a request through the real chain with an element that throws, and asserts the translation point saw it. It fails the day someone adds a broad catch anywhere inside. - A **counter per sanctioned recovery**, so degradation is a number on a dashboard rather than an assumption. - A **review rule** that any new catch in a shared hook either rethrows or comes with the three conditions written down. - **Alerting on the degradation counters**, since the entire point of recovery is that the request still looks fine. ## How it goes wrong The two failure modes are opposites. Catching too much produces a service whose error rate is beautiful and whose behaviour is quietly wrong; the incident arrives weeks later as a data problem rather than an availability one. Catching too little — or placing translation too deep — produces inconsistent responses for the same fault and an error stage whose output nobody trusts. Between them sits the practical rule that survives contact with many teams: **observe freely, recover rarely, and never silently**.

  • What is the cheapest way to stop this policy from eroding?
    A test that drives a request through the assembled chain with an element that throws and asserts the translation point saw it. It fails the day someone adds a broad catch. Pair it with a counter on every sanctioned recovery so silent swallowing becomes visible rather than assumed.
  • When is recovering inside a hook clearly the right call?
    When the concern is genuinely optional and the degraded result is defined — an enrichment lookup that can be skipped, a cache that can be missed. The response is plainer but still correct, and the recovery is counted so a persistent inner failure is not invisible.
  • Why is recovering from a failed permission check treated as a security defect?
    Because the request then proceeds without the check ever having passed, and it proceeds silently. A failure to evaluate a permission is not a permission granted; the safe behaviour is to let it unwind so the translation point refuses the request and the failure is recorded.

saying these in an interview costs you the question

  • Lets every hook catch broadly so that requests never fail outright
  • Recovers from a failed permission check by continuing with a default identity
  • Counts a silent catch as resilience because the error rate went down
  • Adds a second failure-to-response translation deep inside the chain
  • Assumes a cleanup construct is the same thing as handling the failure