A plain text file carries no type information. How does a reader decide each column's type, and why might it decide differently next month?
answer
- nothing in the file says
- one representation per column
- decided from a sample, not a statement
- same bytes, different answer
- state the types at the read
basics
~20 sA reader over a file with no declared types guesses each column's type from the rows it inspects, then holds every value in that column to that one representation. Different rows next month can produce a different guess.
solid answer
~50 sA file that stores every value as characters says nothing about what those characters mean, so the reader has to decide. It inspects some rows, picks one stored representation per column — the single physical form every value in that column is kept as — and applies it to the whole column. The decision is made from a sample, not from the file's own statement of intent, so it is a guess with no authority. Add rows where the pattern breaks, read the file in pieces, or run the same read somewhere configured differently, and the same bytes can come back as a different type. The cure is not to tune the guess but to remove it: state the types when you call the reader, so there is nothing left to infer and a value that contradicts you becomes visible instead of silently accommodated.
go deeper
Remember that a character-only file carries no types, so whatever type your column has was chosen by the reader from a sample of rows. Be able to say that, and that you can state the types instead.
Explain that the choice is one representation for the whole column, made from part of the file, and name the three ways the same file can read differently: new rows, a piece-by-piece read, and different reader settings.
Show that you treat a read as a step with defaults and failure modes. Say how you would make a moved guess loud — types stated at the read, a mismatch failing the step — rather than relying on somebody noticing.
Frame it as who should be doing the declaring. A guess that keeps moving is a signal that the producer is shipping a shape with no types in it, and that may be the cheaper thing to change.
## What the reader is actually doing **The reader** is the call that turns a file into a table in one step: bytes in, a rectangle of typed columns out. For a shape that stores every value as characters — one record per line, fields separated by a delimiter — nothing in those bytes states whether a field is a whole number, a decimal, a date, an identifier or free text. Every field is just characters. Something has to decide what they mean, and that something is the reader. The decision is **per column, not per value**. A column in a labelled table — a rectangle whose columns each carry one type and whose rows may carry an identity of their own — holds one **stored representation**: the single physical form every value in that column is kept as, fixed for the whole column. So the reader is not asking "what is this field?" a million times. It asks "what is this column?" once, from the fields it happened to look at, and then holds every remaining value to that answer. ## When there is a guess at all It is worth being precise, because "the reader infers the types" is only true of one situation: | What the file gives the reader | What the reader does | |---|---| | No type information anywhere (characters only) | Inspects some rows, picks one representation per column, applies it to all of them | | Types written into the file itself | Reads the declaration and infers nothing | | No type information, and the reader is built not to guess | Leaves every column as text until you state otherwise | | No type information, but you stated the types at the read | Uses what you stated; there is nothing to guess | The inference question exists only in the first row of that table. Designs differ here, and the difference is the first thing to establish about any reader you are using. ## Why the same file reads differently Three mechanisms, all ordinary: 1. **The data moved.** If the reader fixes a type from a bounded prefix of the file, then a value that contradicts the pattern matters only if it falls inside that prefix. The same awkward value near the top and near the bottom produces two different reads of the same file. Next month's file has different rows near the top. 2. **The read was cut into pieces.** Where a reader hands back batches of rows rather than the whole table, some designs decide types afresh for each batch. Two pieces of one file can then disagree with each other, and the disagreement only shows up when the pieces are put together. 3. **The read ran somewhere else.** Readers expose settings for how much they inspect and what they do with a contradiction, and those settings are environmental as often as they are deliberate. A job that reads one way on a laptop and another way on a scheduler is not mysterious; it is two configurations of the same guess. None of these is an error condition. No exception is raised, nothing is logged that anybody reads, and the table looks entirely normal. ## What the guess costs when it is wrong - A column of identifiers that happen to be digits becomes numbers, and whatever made them identifiers — leading zeros, length, plus signs — is gone at parse time rather than at print time. - A column the reader decided was text sorts and compares as text, which is a different ordering from the one you meant. - A column's width — how many bytes each value occupies — follows from the representation, so a guess is also a memory decision made on your behalf. - Most importantly: **the failure is silent**. A wrong type does not stop the job, so the first sign of it is a number in a report that nobody can explain. ## Taking the guess away Stating the types when you call the reader costs one statement and changes the failure mode completely. Instead of the reader quietly accommodating whatever it found, it now has a claim from you to check against, and a value that violates the claim is something it can refuse. That is the whole trade: you give up the convenience of not thinking about it, and you get a loud failure at the file boundary instead of a quiet wrong answer three steps downstream. The second route is to read a shape that carries its own type declaration, in which case the reader has nothing to infer. Either way, the thing to internalise as a junior is simple: **when a file has no types in it, the types in your table were invented by the reader, and they are only as stable as the rows it happened to look at.**
- Does a reader over a self-describing file guess in the same way?No. Where the file writes its own type declaration alongside the values, the reader reads that declaration and infers nothing, so none of the instability applies. The guess is a property of reading a shape that carries no types, not a property of readers in general.
- If the guess is unstable, why not simply inspect the whole column every time?Some readers do, and it is more stable for a given file. It costs either a second pass over the file or enough room to hold the column before deciding, which is exactly what a bounded prefix is avoiding. And it still changes answers when the file changes — it removes position-dependence, not dependence on the data.
- The guess was right for two years. Is that evidence it is safe?No. It is evidence that no contradicting value has yet appeared where the reader looks. The guess has the same authority on day 800 as on day one, which is none; the only thing that changed is your confidence in it.
saying these in an interview costs you the question
- Believes a text file stores each column's type somewhere in it
- Assumes the inferred type is stable across runs of the same job
- Thinks the reader checks every value and re-decides as it goes
- Treats a changed type as the data changing rather than the guess moving
- Thinks stating the types up front is more work than it is worth
- Assumes every reader infers types, including over self-describing files