skip to content

A column became text at step two, but the run only fails at step six. Why does the failure surface there?

level: middleimportance: should knowfreq 58%

answer

  1. an error is not what propagates
  2. the value travels, the fault does not
  3. the first step with no meaning on text
  4. strict typing shortens the distance

basics

~20 s

Every step between the change and the failure was defined on text, so each one succeeded and passed the column along. The failing step is simply the first one that required a numeric column and could not get one.

solid answer

~50 s

A representation change does not propagate an error, it propagates a value. Ordering, comparing, grouping, matching and writing are all legal on text, so each intermediate step returned something and handed the column on unchanged. The failing step is the first one in the file with no meaning on text — an arithmetic fold, a division, something that needs a numeric column and nothing else. It names itself in the error, which is how correct code ends up blamed. How far the symptom travels is a property of the design: strict per-column typing refuses at the first step mixing text with a number, so the distance is one step; a catch-all representation can carry the column to the end of the file and, if no step ever demands a number, may never raise at all.

go deeper

for a junior

Recall that the step named in an error is not always the step at fault. A column can change how it is held and then travel through several later steps that are perfectly happy with the new form.

for a middle

Explain the propagation: a value moves from step to step, a fault does not. Name what the failing step needed that text could not provide, and why the steps before it were satisfied.

for a senior

Say out loud that the distance depends on the design — strict per-column typing refuses immediately, a catch-all may never refuse — and that the silent finish is the outcome to plan against.

for a principal

The call is how much distance between cause and symptom the team will tolerate, and what it is worth paying in claims written into the file to shorten it before the file gets long.

## What travels between steps is a value, not a fault When a step changes how a column is held, it produces a column and hands it on. No marker is attached saying this is not what the next step expects, and nothing inspects it in between. Each later step receives a perfectly well-formed column of text and asks one question about it: *can I do my operation on this?* Four times in a row, the answer was yes. So "the run failed at step six" means exactly one thing: **step six is the first step in this file with no defined meaning on text.** It is a lower bound on where the change happened, and nothing more. ## The steps in between were not lucky Ordering compares characters. A comparison against a threshold compares characters. Grouping puts equal spellings together. Matching pairs identical spellings. Writing the table out writes characters. Every one of those has a correct, well-defined answer for text, so every one of them returned a result — a result that is wrong for the question being asked, but a result. That asymmetry is why this class of defect travels: - a step that **cannot** work on text refuses, and stops the run; - a step that **can** work on text succeeds, and stops nothing; - in an ordinary analysis file there are far more of the second kind than the first, and they usually come first. ## Why the failing step gets the blame The error names the operation that refused. The operation that refused is, by definition, the one thing in the file strict enough to notice. The code there is usually correct — a mean, a division, a numeric aggregation — and is the only honest participant in the sequence. A first reaction of "why is this arithmetic broken?" sends the investigation to the wrong place. The useful reading is: *the column arriving here is not what this step was written for; when did it stop being that?* ## The distance depends on the design, and on your file | design | at the first step mixing text with a number | distance between cause and symptom | |---|---|---| | strict per-column typing | refuses outright | short, often a single step | | a catch-all representation holding arbitrary values | succeeds, returning a text-shaped answer | as long as the file — possibly no failure at all | The file matters as much as the tool. A pipeline whose only arithmetic sits in the final summary has nowhere to fail earlier, whatever the tool would have done. A pipeline that multiplies a rate by a quantity in its second step gives the defect nowhere to hide. ## The outcome with no failure at all A run that fails at step six is annoying. A run that never fails is worse, and it is a realistic outcome wherever a catch-all representation meets a file that never demands a number: - ordering, grouping, extremes and counts all return values that exist in the data; - a threshold comparison returns a subset of plausible size; - the table is written out without complaint; - the report is produced, delivered, and believed. The late failure at least converts a wrong answer into a stopped run. Treat it as the better of two bad outcomes, not as the defect itself. ## Shortening the distance on purpose The distance is not a fact of nature; it is a consequence of nothing in between having an opinion. You can give the file one: 1. State beside the step, in the same file, that the column is held numerically at the points where later work depends on it — the run then stops where the claim is false rather than where the arithmetic is. 2. Put those claims at boundaries where the representation can actually change, not after every step: a claim that never fails is a cost with no return. 3. Make the claim about how the column is held, not about whether its values could be read as numbers — a column of digits held as text passes the second and fails the first. The measure of success is not that the run fails less often. It is that when it fails, the step named in the error and the step at fault are the same step.

  • What if the run finishes to the end with no error at all?
    Then the wrong number ships. Where a catch-all representation meets a file that never demands a numeric column, nothing in the run has the standing to refuse: the extremes, the counts and the groups are all real values from the data, so the output looks exactly as it should. That is the outcome the assertions exist for.
  • Does the failing step tell you which step changed the representation?
    No. It tells you where a numeric column was first genuinely required, which is a bound and not a location: the change happened at or before that point, possibly many steps before. Nothing in the message carries the column's history, because no step recorded any.
  • Why is a failure at step six still better than the alternative?
    Because it fails. A stopped run puts the defect in front of somebody while the context is still live, and the cost is a misleading first suspect. The alternative is a finished run whose number nobody questions, discovered weeks later or not at all.

saying these in an interview costs you the question

  • Blames the step named in the error message
  • Assumes the first step touching the column must fail
  • Believes a run that finishes proves the types were right
  • Thinks the distance is the same in every tool
  • Says the intermediate steps must have silently converted the values