skip to content

What does a recovery step that substitutes an empty recommendation list for every failure hide from you?

level: seniorimportance: nice to knowfreq 32%

answer

  1. not every failure is the same failure
  2. the page looks healthy
  3. a defect dressed as an outage
  4. restrict by failure type
  5. translate what you will not substitute

basics

~20 s

Every failure that is not the recommendation dependency being unavailable: defects raised by stages above the recovery step, malformed responses, configuration mistakes. All of them render as a healthy-looking empty block, so the page never breaks and nobody learns it is broken.

solid answer

~40 s

A recovery step matches a failure and substitutes for it; an unrestricted one matches *everything* that reaches it, not just the failure you had in mind. A defect in a mapping stage above it, a malformed response the parser rejected, a misconfigured endpoint — each arrives as a failure and each becomes the same empty list. The page renders, so the symptom is invisible and the only evidence is a block that is always empty. Restricting the recovery by failure type fixes it: the failures the block is genuinely allowed to hide are substituted, and everything else travels on and can still end the sequence. For failures that belong to the caller's decision, translate them into a domain failure rather than inventing data.

code

pseudocode · 11 lines
pseudocode
// substitutes for everything, including a defect above it
block = fetchRecommendations(id)
          .mapEach(r => toCard(r))
          .recoverWithValue(EMPTY_LIST)

// substitutes only what this block is allowed to hide
block = fetchRecommendations(id)
          .mapEach(r => toCard(r))
          .recoverWithValue(EMPTY_LIST,
                            when: failure is DependencyUnavailable)
// any other failure travels on and still ends the sequence

go deeper

for a junior

Recall that recovery can be limited to particular kinds of failure, and that substituting for every failure is a deliberate choice with a cost.

for a middle

Explain what a blanket substitution swallows: defects raised by stages above it and contract breaks, not only an unreachable dependency — and all of them then look like an empty result.

for a senior

Demonstrate the judgment: name the failures the block is allowed to hide, let the rest travel on, and translate transport-level failures into domain failures the caller can act on.

for a principal

Set the boundary between substitution and translation across teams: which layer may invent data, and which must surface a typed failure for its consumers to decide on.

## Substituting for 'any failure' is a wider promise than it sounds A recovery step is defined by what it **matches**. Written without a restriction, it matches every failure that reaches it — and what reaches it is not only the failure you were defending against. On a product page, an unrestricted recovery around the recommendation block will equally swallow: - The dependency being unreachable or timing out, which is the case you wanted to cover. - A **defect** in a stage above it — a bad conversion, a missing field dereferenced, an assumption that did not hold. This is not an outage; it is a bug, and it will never be fixed if it renders as an empty block. - A **malformed response** the parser rejected, which usually means the contract changed and every request is now failing. - A **configuration mistake**, which typically fails one hundred percent of the time and so produces a permanently empty block that looks like a product with no recommendations. - A failure whose correct treatment is to be surfaced to the caller with meaning, not replaced with data. All five become one indistinguishable outcome, and the outcome is *success*. ## The failure mode: a page that is fine The reason this is a senior question and not a style preference is the shape of the resulting incident. There is no failure signal, no failed request, no broken render. There is a block that is empty — and empty is a legitimate state for a recommendation block, so nothing about it is obviously wrong. The two situations below are identical from the outside: | Situation | What the subscriber gets | What is actually true | |---|---|---| | No recommendations exist for this product | Empty list, successful ending | The system is working | | A defect above the recovery step throws on every request | Empty list, successful ending | The block has been broken for weeks | A blanket substitution is the mechanism that collapses those two rows into one. ## Restricting by failure type The correction is to make the recovery step state which failures it is allowed to hide. Concretely, that means a predicate or a type restriction on the recovery, chosen by asking one question per failure: *is substituting an empty list here an honest answer?* 1. **Dependency unavailable** — yes. The block's purpose is optional enrichment, and an empty list is honest. 2. **Defect in our own stages** — no. There is nothing to degrade to; the code is wrong and must surface. 3. **Contract violation in the response** — no. Substituting would mask a break that affects every request. 4. **Cancellation because the caller went away** — no. Nobody is waiting for the fallback; substituting invents work for no one. What the restricted step does *not* match keeps travelling downstream, unchanged, and can still end the sequence — which is exactly what you want, because a failure that ends the sequence is a failure somebody will see. ## Translation is the other half of the answer Between 'substitute data' and 'let it travel on raw' there is a third move: **translate** the failure. A transport-level failure — a connection refused, a timeout — carries no domain meaning, and forcing every caller to interpret transport vocabulary spreads that vocabulary through the system. Translating it into a domain failure at the edge of the block keeps the sequence failed, so nothing is invented, while giving whoever handles it something it can act on. That also settles the placement question for shared code. When the same recommendation source is consumed by three different pages, the source itself should not decide that an empty list is acceptable — one of those pages may not be able to degrade at all. The source fails, translated into a domain failure; each consumer attaches its own recovery and makes its own choice. Substitution is a **consumer's** decision, translation is the **provider's**. ## And the block that may never degrade The same reasoning is what keeps a blanket recovery away from the price lookup. A missing recommendation block is a thinner page. An invented price is a wrong page, and no rendering rescues it. So the rule is not 'add fallbacks generously' but 'name the failures each block may hide, hide exactly those, and let everything else be seen'.

  • What is the difference between translating a failure and substituting for it?
    Substitution ends the failure and supplies data, so the sequence completes normally. Translation keeps the sequence failed but replaces a transport-level failure with a domain-meaningful one, leaving the decision to the caller. Translate when you know what went wrong but not what should be shown instead.
  • Where should the decision to substitute live when three pages consume the same recommendation source?
    At each call site, not inside the shared source. One of those pages may not be allowed to degrade, and a source that substitutes on its own behalf removes the choice from all three. The source should fail — ideally with a translated domain failure — and each consumer attaches its own recovery.

saying these in an interview costs you the question

  • Says a blanket fallback is always safer because the page never breaks.
  • Cannot name a failure that must not be substituted with an empty list.
  • Treats a defect in a stage above the recovery step as a dependency outage.
  • Applies the same fallback rule to the price block as to recommendations.
  • Thinks translating a failure into a domain failure is the same as recovering from it.