skip to content

Before a reader fixes a column's type, what might it have inspected, and how does each choice change the answer?

level: middleimportance: should knowfreq 56%

answer

  1. four rules, not one
  2. prefix, whole column, per piece, refuse
  3. row order can decide the type
  4. pieces can disagree with each other
  5. never quote a row count

basics

~20 s

Readers differ: some fix a type from a bounded prefix of the file, some scan the whole column first, some decide afresh for every batch of rows, and some refuse to guess at all. Which rule applies decides whether a late contradicting value is ever seen.

solid answer

~50 s

There are four rules in circulation and they give genuinely different answers on the same file. A reader that decides from a **bounded prefix** is cheapest and single-pass, but a value that breaks the pattern after the prefix is never seen — so the same value on line 20 and on line 20,000 produces two different reads. A reader that inspects the **whole column** is stable for a given file and pays for it with a second pass or with room to hold the column before deciding. A reader handing back rows in pieces may decide **afresh per piece**, so two pieces of one file can disagree and the disagreement only surfaces when they are combined. And some readers **refuse to guess**, leaving everything as text until you declare. Establish which rule you are under before you trust any inferred type — and never reason from a specific number of inspected rows.

go deeper

for a junior

Know that the reader only looked at part of the file before deciding, and that the part it looked at can decide the answer. That alone explains most surprises here.

for a middle

Name the rules — bounded prefix, whole column, afresh per piece, refuse to guess — and say what each costs and how each fails. Avoid quoting a specific row count; the rule is the content, the number is not.

for a senior

Show that you treat the sampling rule as a property to look up rather than assume, and that your response to an unstable type is to state it rather than to widen the sample.

for a principal

The judgment is when inference is acceptable at all. Exploratory work can live with a guess; anything a report or a downstream job depends on should not have a type decided by which rows happened to be near the top.

## Four rules, not one "The reader looks at the first few rows" is the folk answer, and it is one of four real designs. The rule in force determines what the guess is worth. | What the reader inspected | What it costs to decide | How it goes wrong | |---|---|---| | A bounded prefix of the file | Cheapest: one pass, nothing buffered | A contradicting value after the prefix is never seen, so position in the file decides the type | | The whole column | A second pass over the file, or enough room to hold the column before choosing | Stable for this file, but still moves when the file changes; slower and heavier on every read | | Afresh for each piece of a piece-by-piece read | The prefix cost, paid once per piece | Two pieces of one file disagree with each other, and nothing says so until they meet | | Nothing — the reader declines to guess | Nothing | Every column arrives as text; nothing is destroyed and nothing is convenient | **Never reason from a specific number of inspected rows.** A concrete default identifies a product as plainly as a function name does, and the number is the least stable fact about any of these designs. What matters is the *rule*, not the count. ## Why the bounded prefix is position-dependent This is the property people find most surprising, so it is worth stating baldly. Under a bounded-prefix rule, a file is not read as a whole; it is judged by its opening. Put a value that contradicts the pattern near the top and the reader sees it and chooses accordingly — perhaps a wider representation, perhaps text. Put the identical value near the bottom and the reader never sees it, having already committed. So two files with **exactly the same multiset of values**, differing only in row order, can produce columns of different types. Nothing about that is an error, and nothing warns you. It also means an inferred type is quietly a statement about your data's *sort order*, which is not a property anybody thinks they are depending on. ## Why a piece-by-piece read is a different problem When a reader hands back batches of rows rather than one table, each batch is its own read as far as inference goes, under designs that re-decide per piece. The consequences: - A column can be numeric in the first pieces and text in a later one, because the later piece contains a field the earlier ones did not. - Nothing compares one piece's decision with the next one's, so nothing raises. - The mismatch surfaces later, when the pieces are put together or written out, and it surfaces as a confusing error a long way from its cause. - Reducing each piece to a running total is not immune: two pieces that disagree about a column's type can produce partial results that are not comparable. The cure is not to make the pieces bigger. Enlarging the sample makes the disagreement rarer and no less possible, which is the worst of both worlds: the same defect, now harder to reproduce. ## Why the whole-column scan is not simply better Inspecting the entire column before deciding is the stable answer for a given file, and it has a real price: 1. **Time.** Either the reader passes over the file twice, or it holds the parsed-but-undecided column in memory until the last row is in. 2. **Memory.** The second option is exactly the resident cost you may have been trying to avoid by reading in pieces at all. 3. **It is still a guess.** Stability against row order is not stability against tomorrow's file. A column that is entirely digits today and gains one non-numeric field tomorrow changes type under every one of these rules. ## What to do with this The practical stance is short. First, find out which rule your reader is under — it is documented, and it is the single most useful fact about a reader you are going to depend on. Second, do not respond to instability by tuning how much it inspects; that adjusts the odds rather than the mechanism, and the failure remains silent when it does occur. Third, state the types at the read for every column whose type you actually care about. Once the types are stated, none of this matters: the reader has nothing to infer, the sampling rule becomes irrelevant, and a value that contradicts your statement is something the reader can act on rather than absorb. The short version to carry into an interview: **whether the guess comes from a bounded prefix, from the whole column, or afresh per piece is the thing to establish before trusting it** — and the fact that you have to establish it at all is the argument for not guessing.

  • Two pieces of one file disagreed about a column's type. Why did nothing raise at the time?
    Because each piece was decided independently and nothing holds the earlier decision to compare against. The reader is answering a fresh question per piece, and both answers are locally correct. The contradiction only becomes visible when the pieces are combined or written, well after the read that caused it.
  • Is a reader that scans the whole column before deciding strictly safer?
    Safer against row order, yes — the same values in any order give the same type. Not safer against change: tomorrow's file with one new non-numeric field gets a different type under this rule too. And you pay a second pass or the memory to hold the column, which may be the very cost you were avoiding.
  • Why not just enlarge the sample until the problem stops appearing?
    Because it changes the probability, not the mechanism. The failure is still silent when it happens, and it now happens rarely enough to be hard to reproduce and easy to blame on something else. Stating the types removes the mechanism instead.

saying these in an interview costs you the question

  • States a specific default number of rows as if universal
  • Assumes every reader inspects the whole column before deciding
  • Thinks a larger sample fixes the instability rather than hiding it
  • Believes two pieces of one file must agree on a column's type
  • Says inspecting the whole column is free
  • Thinks row order cannot affect an inferred type