skip to content

A stage maps each search result through an at-most-one lookup and emits fewer values than it received, with no failure - why?

level: seniorimportance: should knowfreq 38%

answer

  1. count what each input contributes
  2. zero or one per input
  3. an empty inner source is a filter
  4. misses are successful endings
  5. default inside the inner source

basics

~20 s

Every inner lookup that completes empty contributes zero values, so n inputs produce between zero and n outputs. Absence is a normal ending rather than a failure, so the misses vanish silently unless the empty case is given a value of its own.

solid answer

~40 s

The arithmetic is the explanation. Each input value is replaced by an inner source declared to carry **at most one** value, so each input contributes zero or one output and n inputs yield 0..n outputs. A lookup that finds nothing ends with an empty completion, which is a normal terminal outcome - not a failure - so nothing is logged and nothing is raised. The count simply shrinks. There are two honest fixes: give the empty case a value inside the inner source, substituting an explicit *not found for this input* marker so every input yields exactly one output; or split the pipeline so misses are emitted or counted on a separate path. If a miss should never happen, convert the inner empty into a failure deliberately, so it is loud rather than invisible.

code

pseudocode · 10 lines
pseudocode
// each identifier becomes an inner source carrying at most one record
found = ids.flattenEach(id -> lookupById(id))
// n identifiers -> 0..n records; every miss silently disappears

// parity restored: give the empty ending a value INSIDE the inner source
outcomes = ids.flattenEach(id ->
    lookupById(id)
        .mapValue(record -> foundFor(id, record))
        .defaultIfEmpty(missingFor(id)))
// n identifiers -> exactly n outcomes, misses carried explicitly

go deeper

for a junior

Remember that a lookup returning nothing is not an error, so a stage built on such lookups can legitimately emit fewer values than it received. Count what each input contributes before assuming a bug.

for a middle

Explain the arithmetic - zero or one output per input - and show where a substitute for the empty case has to be attached for count parity to survive.

for a senior

Diagnose it in production: recognise the clean-run shortfall, rule out transients, decide whether a miss is expected, and instrument inputs against outputs so the gap is measured rather than discovered by a human reading a total.

for a principal

Decide the standard: whether misses are carried as explicit outcomes across the organisation's pipelines, what a miss rate must be reported as, and when an absent record is severe enough to fail a run outright.

## The arithmetic, stated plainly A stage that replaces each incoming value with an inner source and emits the inner values produces an output count equal to the **sum of the inner cardinalities**, not the input count. If every inner source is declared to carry at most one value, each input contributes 0 or 1, so for n inputs: - best case: n outputs, when every lookup found something; - worst case: 0 outputs, when every lookup found nothing; - typical case: n minus the number of misses. There is no rule anywhere in the pipeline that says outputs must match inputs. Count parity is something you build, not something you get. ## Why nothing failed An empty completion is a *successful* ending. The inner lookup was subscribed to, it ran, it consulted the store, and it ended having nothing to report. Everything about that is ordinary, so: - no failure signal travels downstream, and no recovery handling fires; - no metric labelled *errors* moves; - the outer sequence continues happily with its remaining inputs; - the trace shows a healthy request whose result is merely shorter than expected. This is precisely why the defect is discovered late and usually by a human noticing a total. The pipeline is behaving exactly as written; what is wrong is that nobody decided what a miss means. ## What the shape of the fix depends on Ask one question first: **is a miss expected?** 1. **A miss is expected and interesting.** Give the empty case a value *inside the inner source*, before the inner values are merged out. Substituting an explicit marker that says *no record for this input* turns each input into exactly one output and carries the identity of the miss with it. Downstream can then split the two kinds apart, count them, or report them. 2. **A miss is expected and uninteresting.** Leave the drop in place, and make it deliberate - document the shrink, and measure inputs and outputs separately so the ratio is observable rather than surprising. 3. **A miss should be impossible.** Convert the inner empty completion into a failure explicitly, so a broken invariant is loud on the first occurrence instead of silently shortening a result set for months. The critical detail in option 1 is *where* the substitution goes. Supplying a default after the inner values have been merged into the outer sequence does nothing for the miss, because at that point the missing value was never there to be defaulted - the outer sequence only carries values that some inner source actually produced. The default has to be attached to the inner source, where the empty ending exists to be intercepted. ## Inner arity and output count | Inner source's declared arity | Outputs for n inputs | Count parity with inputs | |---|---|---| | Exactly one value | exactly n | preserved | | At most one value | 0 to n | lost on every miss | | Many values | 0 to unbounded | unrelated to n | The table is worth internalising as a review reflex: whenever a stage replaces values with inner sources, read the inner arity and immediately know whether counts survive. ## Symptoms that point here - A batch job reports fewer processed records than it read, with a clean run and no failures. - A response body carries fewer entries than identifiers were requested, and the missing ones differ between runs because the underlying data differs. - A reconciliation total drifts slowly as data is deleted upstream, because deletions turn previously-successful lookups into empty ones. - Retrying the job produces the same shortfall, which rules out a transient fault and points at a structural miss. ## The wider lesson Cardinality composes. Each stage of a pipeline has an arity relationship between what it consumes and what it produces, and a pipeline's end-to-end relationship is the composition of those. A stage that filters is openly lossy and nobody is surprised. A stage that looks up is *covertly* lossy, because the loss is expressed as an empty ending inside an inner source rather than as a visible predicate. Treat any at-most-one inner source as a filter with a hidden condition, and the arithmetic stops being surprising. When designing such a stage, decide on the miss policy at the moment you write it. The decision costs one line - a substituted marker, an explicit failure, or a comment saying the shrink is intended - and not deciding costs a long investigation into a total that nobody can explain.

  • Why does supplying a default after the inner values are merged fail to fix the shortfall?
    Because by then the miss no longer exists to be defaulted. The outer sequence only carries values that some inner source actually emitted, so an absent inner value is simply not represented. The empty ending exists only inside the inner source, which is the only place a substitute can be attached.
  • How would you make the shortfall visible without changing the pipeline's output?
    Measure both ends. Count inputs entering the stage and values leaving it, and publish the difference as a miss count rather than inferring it. A ratio that is expected to sit near one becomes an alertable signal the moment data upstream starts disappearing.
  • What if the inner lookup is declared many-valued instead of at most one?
    Then the output count is unrelated to the input count in both directions: it can shrink on empty inner sources and grow when one input matches several records. Parity is no longer even a meaningful goal, so downstream logic must not use counts to pair results with inputs.

Handing a stack of room numbers to a receptionist and getting back a shorter stack of guests: nobody made a mistake, some rooms were simply empty.

saying these in an interview costs you the question

  • Expects output count to match input count by default
  • Assumes a lookup that finds nothing raises a failure
  • Blames the shortfall on a transient fault and adds retries
  • Supplies a default after the inner values are already merged
  • Uses positional pairing of outputs to inputs after such a stage
  • Concludes the pipeline is broken when nothing failed