skip to content

How would you decide where error boundaries belong across a long-lived application, and which failures would you deliberately leave uncontained?

level: principalimportance: should knowfreq 40%

answer

  1. can it fail independently?
  2. is the page still usable without it
  3. each boundary needs a fallback and owner
  4. never wrap the shell or a critical value
  5. measure fallback rate and reset success

basics

~20 s

Put a boundary where a region can fail on its own and the page is still worth using without it; leave it out where a fallback would strand the user or make a wrong page look complete.

solid answer

~50 s

I treat coverage as a policy with two questions per region. First, **can this fail independently?** Its own data source, an embed, user-supplied content, an infrequently exercised feature — all yes. Second, **is the rest of the page still useful without it?** If yes, it gets its own boundary, a designed fallback, a reset that changes something, and an owning team. I keep one outermost handler as the last resort and one per major view, and I deliberately do *not* wrap two things: the shell, because a fallback there has nowhere to send the user, and anything correctness-critical — a total, a permission state, a safety-relevant value — because a page that looks complete while silently missing that number invites a wrong decision. Then I instrument: fallback rate per boundary, reset success rate, and an alert when a boundary starts firing.

go deeper

for a junior

Take away the placement heuristic: a boundary belongs around a part that can break on its own while the rest of the page stays worth using, and each one needs a fallback someone designed.

for a middle

Be able to argue placement both ways for a concrete page, and to name what each boundary costs beyond the wrapper — fallback, copy, reset path, test.

for a senior

Show that you instrument it: fallback rate per boundary, reset success, uncontained-error rate, and an owner for every boundary so an alert has somewhere to land.

for a principal

Hold the line that containment is wrong where partial output misleads, keep the policy explicit and reviewed per release, and decide first-response failure behaviour per route rather than per component.

## Coverage as a policy, not an instinct Boundary placement is usually decided ad hoc: someone hits a blank page, wraps the culprit, and moves on. That produces a mosaic of undesigned error states and no story for the failures nobody has hit yet. The alternative is a short policy the whole codebase follows, applied region by region. Two questions decide each region: 1. **Can it fail independently?** It has its own data source, it renders user-supplied or third-party content, it embeds something you do not control, or it exercises a rarely used code path. 2. **Is the rest of the page still worth using without it?** If the answer is yes, containment is a real improvement for the user. If no, containment only changes which screen the user is stuck on. Both yes means it earns its own boundary. Only the first means you contain it at the next level up, where a fallback can say something meaningful. ## Choosing the handler shape The three shapes frameworks offer are not interchangeable in practice: - a **boundary component** wrapping a subtree — explicit in the markup, easy to place per region, the default choice for anything that must degrade in place; - a **captured-error lifecycle hook** on an ancestor that already exists — avoids an extra wrapper for a component that is already the owner of that region, and convenient when the fallback is that component's own alternate state; - an **application-level handler** registered with the runtime — the last resort and the natural reporting funnel, but it has the whole application as its granularity. A workable default: application-level handler for reporting plus the final backstop; a boundary component per view and per independently failing widget; a lifecycle hook where the owning component is already there and a wrapper would be noise. ## What I deliberately leave uncontained | Region | Why no boundary of its own | |---|---| | The application shell and navigation | the fallback would have nowhere to send the user; a broken shell is a crash and should read as one | | A correctness-critical value — a total, an entitlement, a safety indication | a page that looks complete while quietly missing it invites a decision on wrong information | | A whole-page layout a single boundary already wraps | extra nesting buys nothing when the inner region cannot be lost independently | | Anything whose failure means the application's own state is suspect | continuing is the risk; offering a reload is the honest move | The general rule: **containment is right when partial output is honest, and wrong when partial output is misleading.** A dashboard missing one chart is honest if it says so. An invoice missing its total is not, no matter how politely the placeholder is worded. ## The costs nobody budgets for Each boundary is not one wrapper. It is: - a fallback design and its copy, in every supported language; - a reset path that changes something, plus an attempt cap; - a decision about what the fallback may reveal about the failure; - a test that the fallback renders and that the reset works; - an owner who is told when it fires. That is why "a boundary around every component" fails as a strategy: the mechanism scales, the design work does not. Coverage should track the regions the product genuinely wants to keep alive independently. ## Server-rendered and first-load failures A throw while the first response is being produced has no mounted tree to swap a fallback into, so the choice is different in kind: degrade the response (send a shell and let the client attempt the region again), or fail the request with an error status. Decide that per route rather than per component, and make the decision explicit — an uncaught server-side throw that becomes an empty page is the worst of both. ## Making the policy visible A containment strategy that is not measured decays. The instrumentation I want: 1. **Fallback rate per boundary**, per release. A boundary that starts firing is a regression signal with a location attached. 2. **Reset success rate.** A reset that never succeeds means the fallback is a dead end and the boundary is only a nicer crash. 3. **Uncontained-error rate.** Throws that reached the outermost handler mark the places the policy has not covered yet. 4. **Ownership.** Every boundary maps to a team, so the alert lands somewhere. ## How I would explain the tradeoff to a team Containment converts a loud failure into a quiet one. That is a gain when the quiet failure leaves the user able to work and the report still reaches engineering; it is a loss when it leaves them stranded, or when it lets a page assert something false. So the standard I hold is: **every boundary owes the user an action and engineering a report.** A boundary that provides neither should be removed, and the failure allowed to be loud again.

  • Why not simply wrap every component and be safe?
    Because the mechanism scales but the design work does not. Every boundary needs a fallback, copy, a reset that changes something, a test and an owner. Wrapping everything yields hundreds of undesigned error states, a page that can degrade into an unreadable mosaic, and no signal about which failures matter.
  • How do you justify leaving a correctness-critical region uncontained?
    Because the failure mode of containment there is a page that looks complete and is wrong, which a user can act on — approving an amount, trusting a permission state. A loud failure forces the user to stop. If the region must degrade, the degraded state has to be unmistakably marked as unknown rather than rendered as a value.
  • What tells you the policy needs revisiting after a release?
    A rise in throws reaching the outermost handler points at regions the policy never covered. A boundary whose fallback rate jumps is a located regression. And a boundary with a high fallback rate but a low reset success rate is one whose recovery path does not work, which is a design defect rather than a coverage gap.

saying these in an interview costs you the question

  • Proposes a boundary around every component as a safety strategy
  • Wraps the shell and navigation, leaving the fallback with no escape route
  • Contains a correctness-critical value behind a placeholder that reads like data
  • Counts boundaries declared as the coverage metric instead of failures contained
  • Ships boundaries with no owner, so nobody hears when one starts firing