skip to content

Your teams keep losing asynchronous failures in production streams — what visibility standard would you set, and how would you make it stick?

level: principalimportance: should knowfreq 38%

answer

  1. make the safe path the easy path
  2. standardise outcomes, not tooling
  3. an undeliverable signal is a defect
  4. configure the sink at startup
  5. enforcement mechanism is the design

basics

~20 s

Set a small outcome contract every stream job must meet, then make the compliant path the easiest one to write rather than a rule to remember. Standardise what must be observable, never which library produces it.

solid answer

~40 s

Standardise outcomes, not tooling. Four requirements cover almost all of the loss: every terminal subscription registers a failure handler; the process-wide sink for undeliverable signals is configured at startup to count and log with a stack; every run records one terminal outcome — completed, failed or cancelled, with a reason; and every failure carries the work item's identity plus a correlation identifier from the moment it is created. Enforcement is the hard half: ship one sanctioned subscribe helper that does the first three, assert at startup that the sink is configured, and fail the build in test rather than paging in production. Then decide the two real trade-offs: whether an undeliverable-signal event pages or is only counted, and how much diagnostic tracing the platform buys by default versus samples during an incident.

go deeper

for a junior

The takeaway to carry forward: an asynchronous job needs someone listening for its failure and a record of how it ended, or it can die unnoticed.

for a middle

Be able to state the small contract — a failure handler on every subscription, a configured last-resort sink, one terminal outcome per run, identity inside the failure — and why each one closes a specific way failures are lost.

for a senior

Show that you would implement it as defaults rather than rules: a sanctioned subscribe helper, a startup assertion, an automated check before merge, and a burn-down of existing events before any alert is armed.

for a principal

Own the trade-offs explicitly: fail fast in test but never in production on shared workers, price always-on tracing against sampling, and accept that a central helper buys enforcement at the cost of team autonomy.

## Standardise the outcome, not the library The failure mode is organisational, not technical: every team knows how to log an error, and failures still vanish. They vanish because nothing in the codebase makes losing one *harder* than losing one. So the standard should describe **what must be observable about every asynchronous run**, and say nothing about which reactive library a team picked — otherwise it stops applying the moment a team makes a different choice, and it reads as a tooling mandate rather than a contract. ## The contract worth mandating Four requirements, deliberately few enough to remember: 1. **Every terminal subscription registers a failure handler.** A subscription with only a value handler is the single largest source of lost failures, and it is mechanically detectable. 2. **The process-wide sink for undeliverable signals is configured at startup.** It counts, logs with whatever stack exists, and tags the event as its own defect class — never a generic error. 3. **Every run records exactly one terminal outcome** — completed, failed or cancelled, with a reason on the last two — so an ending is a positive fact rather than an absence somebody must notice. 4. **Every failure carries identity.** The work item and a correlation identifier go into the failure when it is created, because nothing derived from a thread survives a hop. Two more are worth *recommending* rather than mandating, because their cost is real: named markers at chain boundaries, and the assembly-time diagnostic mode under sampling. ## Enforcement, in order of how well it works | Mechanism | Strength | Cost | |---|---|---| | One sanctioned subscribe helper that supplies the defaults | Strongest — compliance is the default | Building and owning a shared component | | Startup assertion that the sink is configured | Strong for one requirement, fails loudly | Trivial | | Automated check that flags handler-less subscriptions | Good, catches the main case pre-merge | Rule authoring and false positives | | Review checklist | Weak — decays within a quarter | Reviewer time on every change | | A written policy alone | Effectively none | Nothing, which is the problem | The ordering is the point of the answer. A standard with no enforcement path is a preference, and it will be relitigated by every team that finds it inconvenient. ## The trade-offs a lead actually owns - **Should an undeliverable-signal event fail the process?** In test, yes — failing fast is how a defect becomes visible before it ships. In production, no: those workers are shared with unrelated work, and killing the process converts one team's bug into everybody's outage. Count it, page on a threshold, and keep the loud behaviour where it is safe. - **Page or merely count?** An undeliverable failure is a defect, so it should page while the count is low — which requires the count to *be* low first. Rolling it out into a codebase with hundreds of pre-existing events means counting and burning them down before the page is switched on, or nobody will ever trust the alert. - **How much tracing does the platform buy by default?** Always-on assembly-time tracing captures a stack per stage per subscription and is not affordable on a hot path. Buy markers and carried identity as the default; buy deep tracing as a targeted, time-boxed capability for incidents. - **Central helper or team autonomy?** A shared helper is the strongest enforcement available and the most contentious thing on the list. Keep its surface tiny and its escape hatch explicit, or teams will fork it and you will lose the enforcement you were buying. ## How to roll it out without a mutiny 1. Instrument first and mandate second: turn on the sink's counter everywhere, and publish the number of events per team. The argument usually ends there. 2. Ship the helper before the rule, so the rule is satisfiable the day it lands. 3. Fail the build in test environments only, for one quarter, then tighten. 4. Review the alert's precision after a month. An alert that pages on noise is worse than no standard at all, because it teaches people to ignore the class of defect you were trying to surface. ## What separates this from the senior version A senior answer diagnoses one lost failure and fixes it. This answer is about a contract other teams must live with: it accepts that the mandate's *enforcement mechanism* is the design, not the list of requirements, and it prices the two things that are genuinely expensive — always-on tracing and alert noise — instead of asking for everything and getting nothing.

  • Should the sink for undeliverable signals be allowed to terminate the process?
    In test environments, yes — failing loudly is how a lost failure is found before it ships. In production, no: the workers involved are shared with unrelated work, so terminating converts one team's defect into a service-wide outage. Count the event, log it with a stack, and page on a threshold instead.
  • How do you stop this standard from turning into alert noise?
    Separate the defect class from expected failures: an undeliverable-signal event means nobody took responsibility for a failure, while a handled failure is normal operation. Count both, page only on the first, and burn the existing backlog down before enabling the page. Review the alert's precision after a month and adjust.
  • Why mandate the outcome contract rather than a specific instrumentation library?
    Because the contract survives a team choosing different tooling, a runtime migration, or a new stack, while a library mandate expires with the library and invites arguments about the tool instead of the outcome. Requirements phrased as what must be observable are also checkable at review and in automated rules.

saying these in an interview costs you the question

  • Writes a standard with no enforcement mechanism behind it
  • Mandates always-on diagnostic tracing without pricing it
  • Treats undeliverable-signal events as noise to be filtered out
  • Makes the production sink terminate the process on any event
  • Standardises on one library instead of the observable outcome
  • Assumes more logging by itself produces more visibility