skip to content

Why should a stream run that ends by cancellation be counted as an outcome distinct from one that ends by failure?

level: middleimportance: nice to knowfreq 31%

answer

  1. not every ending is a break
  2. the consumer stopped asking
  3. cancellation travels upstream
  4. three outcomes, not two
  5. record the reason as well

basics

~10 s

Cancellation means the consumer stopped wanting values, not that anything broke. Folding it into the failure counter inflates the error rate during ordinary disconnects; folding it into success hides work that stopped half done.

solid answer

~40 s

A run can end three ways: it completes, it fails, or the consumer cancels. Cancellation is a signal sent upstream by the consumer — a deadline elapsed, the caller went away, a newer input superseded this run, the service is shutting down — and it says nothing about whether the work was broken. Count it as a failure and a burst of client disconnects looks like a failing dependency, which is how teams chase a phantom incident. Count it as a success and you lose the fact that a document was fetched but never indexed. Keep three terminal counters and record a reason alongside the cancellation, because the reasons imply different actions. Expect cancellation counts to move together with undeliverable-failure counts, since cancelling releases the subscriber while work is still in flight.

code

pseudocode · 5 lines
pseudocode
indexing_run
  .on_complete(()      -> counter("run.completed").increment())
  .on_failure(f        -> counter("run.failed", cause(f)).increment())
  .on_cancel(reason    -> counter("run.cancelled", reason).increment())
  .subscribe(on_value = write_to_index, on_failure = log)

go deeper

for a junior

Remember that a run can end three ways, not two: it finishes, it fails, or the consumer cancels. Cancellation means somebody stopped wanting the result.

for a middle

Explain the mechanics: cancellation is a signal the consumer sends upstream, it is not a defect, and merging it into either the failure or the success counter distorts a different metric.

for a senior

Show the operating consequence: record a reason with every cancellation, expect undeliverable-failure counts to move with it, and read a deadline-driven cancellation spike as a latency investigation rather than an error one.

for a principal

Decide the outcome taxonomy other teams inherit — which endings are counted, which reasons are mandatory dimensions, and which of them are allowed to consume the error budget.

## Three endings, not two Every subscription ends exactly once, in one of three ways: 1. **Completion** — the source said it had no more values and the run finished normally. 2. **Failure** — something went wrong and a failure signal terminated the sequence. 3. **Cancellation** — the **consumer** stopped wanting values and said so upstream. The third is the one instrumentation routinely forgets, because most dashboards were designed around a request/response world where there are only two outcomes. The direction matters and is easy to state backwards: values travel **down** the chain, while cancellation travels **up** it, from the consumer toward the source, telling every stage to stop producing and release what it holds. ## Why cancellation is not a failure The ordinary causes of cancellation are all *decisions*, not defects: - A deadline or timeout elapsed and the consumer gave up. - The caller disconnected and nobody is waiting for the result. - A stage that switches to the newest input dropped an older in-flight run in favour of a fresher one — that is the intended behaviour of such a stage, and it cancels the run it abandoned. - The service is shutting down and is tearing subscriptions down deliberately. None of those means the pipeline is broken. If they land in the failure counter, then a burst of client disconnects — a mobile network dropping, a page closed, a batch job stopped by an operator — reads exactly like a downstream dependency failing. Teams have paged themselves at three in the morning over this. ## Why it is not a success either The opposite mistake hides real information. A cancelled indexing run is a document that was fetched, possibly parsed, and never written. The run consumed capacity and produced nothing. If cancellations are counted as successes, the indexing dashboard shows full throughput while the index quietly falls behind — and 'documents indexed' will not match 'runs completed'. | Counting choice | What it distorts | Symptom you will chase | |---|---|---| | Cancellation counted as failure | Error rate, error budget, alerts | A phantom dependency incident during a disconnect burst | | Cancellation counted as success | Throughput and completion rate | An index that falls behind while the dashboard looks fine | | Cancellation counted separately, no reason recorded | Nothing, but the counter is unactionable | Knowing it happened and not why | ## Record the reason, not just the count A bare cancellation counter tells you a number. The reason tells you what to do: - **Deadline elapsed** — look at latency upstream, or at whether the deadline is too tight. - **Consumer went away** — usually normal; alert only on the shape of the curve. - **Superseded by newer input** — expected for a stage that keeps only the latest, and worth watching only if the rate suggests wasted work. - **Shutdown** — expected during a deploy, and a spike outside a deploy window is the interesting case. A hook that observes cancellation gives you the same guarantee as a failure hook: it records the ending and changes nothing about it. ## The interaction worth knowing Cancelling releases the subscriber immediately, but upstream work that is already in flight does not stop instantaneously — a request already sent still returns or times out. If that in-flight work then fails, the subscriber is already gone and the failure has nowhere to be delivered, so it takes the last-resort path for undeliverable signals. This is why cancellation counts and undeliverable-failure counts rise together, and why a spike in the latter during a disconnect burst is usually not a new defect. ## In the indexing job The job runs per document with a per-run deadline. On a bad afternoon, the source it fetches from slows down; deadlines start elapsing; runs are cancelled. With one 'errors' counter, the graph looks like the indexer broke. With three counters and a reason, the graph says: nothing failed, 40% of runs hit their deadline, the fetch is slow — a completely different investigation, started from the first glance at the dashboard.

  • Which cancellation reasons are worth separating in the counter, and why?
    At least four: a deadline elapsed, the consumer disconnected, a newer input superseded the run, and shutdown. Each implies a different action — tune latency or the deadline, watch the curve's shape, measure wasted work, or check whether a deploy was in progress. A single undifferentiated total tells you something happened and nothing about what to do.
  • Why do cancellation counts and undeliverable-failure counts often rise together?
    Cancellation releases the subscriber at once, but work already in flight upstream finishes on its own schedule. If that work then fails, there is no subscriber left to deliver to, so the failure takes the last-resort path for undeliverable signals. A spike there during a disconnect burst is usually this interaction, not a new defect.

A parcel refused at the door and a parcel lost in transit both end the delivery, but only one of them means something is broken — and a courier who files them in the same column will spend the week investigating the wrong thing.

saying these in an interview costs you the question

  • Counts every non-completion as a failure
  • Treats cancellation as success because nothing went wrong
  • Says cancellation is a failure signal travelling downstream
  • Records a bare cancellation total with no reason attached
  • Assumes a cancelled run leaves no half-finished work behind