A read of a file whose tail was never written finishes without complaint — what did the reader hand back?
answer
- nothing in the bytes says how many
- a prefix reads as a complete file
- the last line is the interesting one
- the count must come from outside the file
basics
~20 sA shorter table, with nothing to say so. Delimited text carries no record count, so a file cut off mid-write reads like a shorter complete file; the last partial line either parses wrongly or meets the malformed-line disposition.
solid answer
~50 sOne record per line, every value stored as characters: that shape carries no statement of how many records it should hold, so almost any prefix of it is itself a well-shaped file, up to its last line. The reader has nothing to object to. The interesting part is that final line. If the cut landed where the field count still matches, it parses and you get a record whose values look real and are truncated. If it landed short of that, the malformed-line disposition takes over — stop, drop, or restore the shape. A shape that carries its own declaration of what it contains behaves differently: a cut usually destroys structure the reader must locate first, so the open fails outright, which is the better failure. The defence is the same arithmetic either way: a record count from outside the file, plus a look at the last row.
go deeper
Understand that a file can simply stop in the middle and the read will still succeed. Nothing inside plain delimited text states how many records it was meant to contain.
Explain the fates of the final partial line and why none of them is guaranteed to raise, and say why a shape carrying its own declaration tends to fail more usefully than one that does not.
Show that you check the last row and an externally known record count on every read of a file you did not write yourself, and that you do it before the first aggregate is computed.
Weigh the standing cost of tolerating files you cannot verify against the cost of requiring the producing side to state how many records it sent, and decide which side of that boundary the check belongs on.
## Why a cut file is indistinguishable from a short one A reader over delimited text is handed a sequence of bytes and asked to make records of it. Nothing in those bytes says how many records there were supposed to be. There is no count at the start, no total at the end, no marker that means "this is the end and it was intended to be." So a file that stopped halfway through being written is, from the reader's point of view, **a complete file that happens to be smaller**. Almost any prefix of a well-shaped file is itself a well-shaped file, up to whatever is left of its last line. This is why the failure is so durable. The read succeeds. The table opens. The first rows look right. The row count is plausible. The aggregate is a few percent low, which is inside the range the number moves in anyway. Nothing anywhere is red. ## The three fates of the final line 1. **The cut landed on a boundary.** The last complete record ends the file cleanly and nothing partial survives. The read is perfect and the table is simply missing everything after that point. 2. **The cut landed where the field count still matches.** The line parses. You get a record whose last value is truncated — a name cut in half, a number missing its final digits, a timestamp missing its seconds. It looks like data because it is data, with an ending removed. 3. **The cut left too few fields.** Now the reader has a line that does not fit the shape, and its malformed-line disposition decides: stop, drop the record, or pad the fields it did not find with the absent-value marker, the placeholder a tool puts in a cell with no value. Only the first fate of the third case raises anything, and only where that is the configured disposition. The other paths all return a table. ## What the shape can and cannot notice | Shape | What a truncation removes | What the read tends to do | |---|---|---| | One record per line, values as characters | the tail of the last line, and everything after it | succeeds, returning a shorter table | | Values stored column by column with the declaration written into the file | structure the reader must locate before it can return anything | usually fails to open at all | The second row is the better failure and it is worth naming as such in an interview: a shape that cannot be read at all when it is incomplete has told you something, whereas a shape that reads short has told you nothing. That is a real argument in favour of one shape over another for files you receive rather than write. ## What actually defends against it - **A record count established outside the file.** This is the only control that can contradict a file which is internally consistent but short. The file cannot supply it, because the whole problem is that the file is silent about its own length. - **Look at the last row, every time.** It costs nothing and it catches fate 2 whenever the truncation happens to fall somewhere visible — a name ending mid-word, a timestamp with a missing tail. - **Set the reader to stop rather than pad.** Fate 3 then becomes a failure rather than one silently wrong row at the bottom of the table. - **Do the comparison before the first aggregate.** Once a number has been produced from the short table, the short table has an answer to point at, and nobody re-opens the question. ## Why the eye test fails here specifically Every instinct people use instead of counting fails against this particular failure: - **No error** is expected, because there was nothing for the reader to object to. - **A plausible head** is guaranteed, because the beginning of the file is intact by construction — truncation removes the end. - **A plausible aggregate** is likely, because a missing tail moves most totals by less than they move week to week. - **A plausible file size** is likely too, because sizes vary for innocent reasons and a truncated file lands comfortably inside the usual range. The entire class of failure is characterised by looking normal, which is exactly why the defence is a number established before the read rather than an impression formed after it.
- Why does a self-describing shape usually fail on truncation while plain text does not?Because it carries structure a reader must locate before it can return anything, and cutting the file removes or corrupts part of that structure, so the read fails instead of succeeding short. Plain text has no such structure: it is a sequence of lines, and a shorter sequence of lines is still a file.
- The last row's values all look ordinary. Does that clear the read?No. A cut landing after a separator leaves genuine values behind it, and a truncated one can still read as a plausible name or number. Looking at the last row catches the obvious cases and nothing else; only an externally known record count settles the question.
- What is the cheapest thing that makes this class of failure loud?A record count established outside the file, compared against the table's row count in the same breath as the read. Everything else — the absence of an error, a plausible first page, a sensible total, a normal file size — is entirely compatible with a file that stopped early.
saying these in an interview costs you the question
- Assumes a truncated file cannot be read at all
- Expects the reader to detect an incomplete file
- Never looks at the final row after a read
- Thinks a plausible aggregate proves a complete read
- Believes every file shape fails loudly when cut short