skip to content

Column types are stated only at the first step of a five-step pipeline and a bad value entered at step three — where else should they be stated?

level: seniorimportance: should knowfreq 50%

answer

  1. every construction is a boundary
  2. a step's output is a fresh table
  3. shorten the distance to discovery
  4. refuse where everything should parse
  5. a stale declaration blocks good data

basics

~20 s

At every point a table is built, not only the first. A step that builds a new table from a previous step's output is a fresh construction with no memory of your declaration, so its columns get their representation from whatever values arrive.

solid answer

~40 s

Teams state types where data enters and nowhere else, because that is the boundary that feels like a boundary. But a **construction boundary** is any point a table is built — from records in memory, from another step's output, from a combination of tables — and each one chooses representations again. Step three built a table and stated nothing, so its columns were decided by the values it happened to receive that day, and the bad value was absorbed instead of refused. The fix is to restate the declared column type at each construction boundary, and to prefer a refusing conversion at boundaries where every value is supposed to parse. The cost is honest: a declaration per boundary to maintain, and a stale one refuses good data.

go deeper

for a junior

Know that stating a column's type applies to the table being built right there, and that a later step building its own table is a separate decision. That alone explains most of this failure.

for a middle

Explain what counts as a construction boundary: records in memory, another step's output, a combination of two tables. Say what happens at a boundary with no statement — the representation comes from whatever values arrived.

for a senior

Show the diagnostic instinct: find the boundary nearest the bad value and ask what would have refused it there. Choose a refusing conversion where every value should parse, and be explicit that declarations cost maintenance.

for a principal

The standing decision is how many boundaries your team declares at and who keeps them current. Weigh the friction of stale declarations refusing valid changes against the cost of a figure that was wrong for weeks.

## Where the boundaries actually are Almost everyone puts a type declaration at the point where data first enters the pipeline, because that point looks like a boundary: it is the edge of the system, it has a name, and it is where the mistakes are legendary. Then they write four more steps, each of which builds a new table, and state nothing at any of them. A **construction boundary** is any point at which a table comes into existence. Inside a pipeline these are easy to miss because none of them looks like an edge: - a table built from records already in memory — a collection you assembled, or a result handed back from an in-process call; - a table built from a previous step's output, which is the one in this scenario and the one that is missed most often; - a table produced by combining two tables, where each column's representation is resolved from both inputs; - a table rebuilt after a reshape or a re-ordering, where a step that looks like it only moves data is actually producing something new. At each of these, a column's representation is decided again. If you stated nothing, it is decided from the values present at that moment — and the values present at that moment are the output of your own earlier code, which you were trusting. ## Why step three is the classic place for a bad value to get in The declaration at step one did its job: the data that entered was checked against what you said, and what did not conform was dealt with there. Then step two produced a result, and step three built a table from it. That construction had no declaration, so it took its types from the data. A value that would have been refused at step one was, at step three, simply absorbed — the column's representation adjusted to accommodate it, or the column fell back to a form that holds anything, and nothing was reported. The damage is that the pipeline is now carrying a column whose representation nobody chose, downstream of the only place anyone is looking. The report is wrong for weeks because every subsequent step behaves perfectly on the data it was given. ## The rule, and what it costs The rule is short: **state the declared column type wherever a table is built, not only where data enters.** Three notes on applying it: 1. **Put the statement next to the code that builds the table.** A declaration that lives in a separate place drifts, and a drifted declaration is worse than none where the design enforces it, because it refuses good data. 2. **Pick the conversion disposition per boundary.** At a boundary where every value is supposed to parse — which is true of anything produced by your own earlier step — a refusing conversion is right, because a value that will not convert means an earlier step is broken and you want the run to stop while the evidence is intact. At a boundary receiving genuinely messy input, coercing may be right instead. 3. **Do not declare what you do not know.** A boundary where you cannot honestly say what the type should be is a boundary where a declaration is a guess that will eventually refuse legitimate data. State the columns you are sure about and leave the rest. The cost is real and worth saying out loud in an interview: you now have a declaration to maintain at every boundary, five places to update when a column is added, and a genuine risk that a stale declaration blocks a valid change. The benefit is that the distance between a bad value and its discovery is at most one step. ## How much this matters depends on the design | | Declaration enforced at construction | Declaration is a starting hint | |---|---|---| | One statement at step one | protects downstream, until a new table is built | protects that construction only | | Restating at every boundary | redundant but harmless, and cheap insurance | the only way to get the type you want at all | | What a missing statement at step three does | lets the boundary decide from data | lets the boundary decide from data | The last row is the point. The two designs differ in how much a single statement is worth, and they agree completely about what an *absent* statement is worth at a new construction: nothing. So the rule holds either way, and the design only changes how redundant it feels. ## What an interviewer is listening for Three things. That you can name construction boundaries other than the obvious one, and in particular that a table built from another step's output is one. That you frame the goal as shortening the distance between a bad value and the line that notices it, rather than as declaring types for their own sake. And that you say the cost — declarations to maintain, and a stale declaration refusing valid data — instead of presenting the rule as free. A candidate who has actually run a pipeline like this will also say that they would start with the one or two boundaries where the failure hurt, rather than declaring everything on day one.

  • Why is a table built from another step's output the boundary people miss?
    Because it does not look like an edge. The data came from your own code, one function call away, so it feels already trusted and already typed. But construction is construction: the new table chooses representations from the values it receives, with no memory of what was declared upstream. The feeling of continuity is the illusion, and it is exactly why a bad value entering there travels furthest.
  • Rather than declaring at all five boundaries at once, where would you start?
    At the boundary immediately after the step that produced the bad value, because that is where the evidence is and where a refusal would have cost the least. Then work outward only as far as the failures justify. Declaring everything on day one produces five declarations nobody understands the reason for, and the first time one refuses a legitimate change, somebody deletes them all.

saying these in an interview costs you the question

  • Thinks stating types at the entry point covers the pipeline
  • Does not recognise a step's output as a new construction
  • Presents declaring everything as free of maintenance cost
  • Keeps declarations far from the code that builds the table
  • Declares a type at a boundary they cannot honestly specify
  • Uses one conversion disposition for every boundary alike