skip to content

Two File Shapes

One shape is readable by anything and carries no types; the other reads one column without touching the rest. People choose between them on file size, the difference that decides least.

on this pageshow

questions

4

The same dataset is stored as delimited text and as a self-describing binary columnar file: what does each read have to decide?

level: juniorimportance: must knowfreq 62%

answer

  1. only one shape says what it holds
  2. characters against declared types
  3. the decision lives outside the file
  4. same file, two readers, two answers

basics

~20 s

A self-describing binary columnar file carries its own type declaration, so the read applies it. Delimited text carries characters only, so something outside the file decides what each field means — on every read, and possibly differently the next time.

solid answer

~60 s

Both shapes can hold the same flat rectangle of rows and columns, but only one of them says what it holds. **Re-parsed delimited text** — one record per line, every value stored as characters — carries no types at all, so the decision about what each field means is made outside the file, on every single read. **A self-describing binary columnar file** — values stored column by column in their declared types, with the declaration written into the file itself — is read by applying stated types rather than by establishing them. What that buys is less about parsing work than about stability: two readers, two tool versions or two different days can turn the same text file into differently typed columns, and nothing in the file contradicts any of them. Designs differ in how a text read decides — some sample rows and guess, some hand back every column as characters until you declare — but in all of them the decision lives in the reading code, not in the data.

go deeper

for a junior

Recall the one-line difference: characters with no declaration, against values stored in declared types with the declaration inside the file. Then say what the next read therefore has to do in each case.

for a middle

Explain where the typing decision lives. For the text shape it is in code, not in data: the reader samples and guesses, or returns characters until told. That is why the result is reader-dependent.

for a senior

Show that you treat a column's type as something either guaranteed by the file or established by your code, and that you know which one you are relying on for every input you depend on.

for a principal

Weigh the friction of making a producer declare types against the standing cost of every consumer re-deciding them. The second is paid on every read by everyone, and stays invisible until it is wrong.

## The two shapes **Re-parsed delimited text** means one record per line, values written as characters, with separators between the fields. **A self-describing binary columnar file** means values stored column by column in the types they were declared to have, with that declaration written into the file itself. Both can hold exactly the same flat rectangle of rows and columns, and either can be produced from the other. The difference this question is about is not how large they are on disk: it is how much the *next process to open the file* has to work out for itself. ## What a characters-only file leaves undecided Nothing in a delimited text file says that the third field is a whole number, the fourth a moment in time and the fifth a piece of text that merely looks numeric. Every field is characters. So the decision has to be made somewhere outside the file, and that somewhere is the code doing the reading. Designs differ in how they go about it: - **Sample and guess.** The reader inspects some rows — a prefix, a block, or the whole column — and fixes a type from what it saw. - **Return characters and wait.** Some readers decide nothing at all, handing back every column as characters until you say otherwise. - **Take a declaration you supply.** You state the types up front and the reader has nothing to guess. All three put the decision in the reading code rather than in the data, and that is the point. Two consequences follow directly: 1. **Two consumers can disagree.** Your read and a colleague's read of the same bytes can produce differently typed columns, and neither file nor reader is in a position to say which is wrong. 2. **One consumer can disagree with itself.** A different tool version, a different declaration, or simply a different set of rows in tomorrow's file can move a column's type without anything in the file changing its meaning. ## What a declared file removes When the types are written into the file, the read applies them. The question of inference does not arise for that shape — there is nothing to establish, because the file already states it, and the statement travels with the bytes to everyone who opens them. That is the whole of what "self-describing" means here, and it is why the two shapes behave so differently at a handoff. | Question at read time | Delimited text | Self-describing columnar | |---|---|---| | What is this field's type? | decided by the reading code | stated in the file | | Who holds that decision? | every consumer, separately | the file, once | | Can two consumers disagree? | yes, and silently | not about the declared types | | What does obtaining a value involve? | converting characters, every read | reading it in its stored form | | What travels with the bytes? | the characters only | the characters' meaning too | ## What a declaration does not promise It is worth being precise about the size of the claim. A declaration is a statement about **representation**, not about truth or quality. A file can declare a column as text and fill it with nonsense; declaring a type does not make the values correct, complete or meaningful. Nor does a declaration make the two shapes interchangeable in every respect — text opens in anything at all, including a plain viewer and the eyes of a person in the middle of an incident, while a binary shape needs something that understands its layout. That is a real property the text shape has and the other does not. ## Where the difference actually bites The cost of re-deciding is invisible while one person reads their own file with their own code on their own machine. It shows up at the seams: - **Between people.** Handing someone a text file hands them the job of deciding what is in it, and they will decide with different code than you did. - **Between tools.** A second tool reading the same file brings its own defaults, which were designed for a different balance of convenience and strictness. - **Across time.** The same reading code applied to next month's file can settle on a different type, because the decision was derived from the data rather than declared alongside it. The practical move is to stop treating "what type did this column come back as" as a detail of the read and start treating it as something that is either guaranteed by the file or established by your code — and to know, for any dataset you depend on, which of the two you are relying on. A team that cannot answer that question for its own inputs is one awkward file away from a wrong number that raises nothing at all.

  • If you can hand the reader a declaration, why is that not as good as one inside the file?
    Because it lives with the code rather than with the data. Every consumer needs its own copy, the copies drift, and anyone handed the file without your code is back to deciding for themselves. A declaration inside the file travels with the bytes and reads the same for everyone.
  • Does a read of a self-describing file infer anything at all?
    Not about the column types — it reads them, so the inference question does not arise for that shape. For a characters-only file there is always something to establish, whether by sampling rows, by returning everything as characters, or by taking the declaration you supply.

saying these in an interview costs you the question

  • Claims the difference between the shapes is mainly file size.
  • Says a text file 'has types' because a tool displayed numbers.
  • Assumes every reader guesses types; some return characters and wait.
  • Expects two tools to read one text file into identical types.
  • Treats a declaration you pass the reader as living in the file.
open as a page

A read asks for 3 of a file's 200 columns: what does that avoid on a columnar layout, and what on delimited text?

level: middleimportance: should knowfreq 58%

basics

~20 s

On a declared columnar layout the unread columns are never read from disk at all. On delimited text every byte of every line still crosses the parser to find the separators; only converting and keeping the unwanted columns is avoided.

open as a page

A process must keep adding records to a dataset on disk: what does appending cost on each of the two file shapes, and what accumulates?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Delimited text takes new records by concatenation, so appending is nearly free. A finalised typed file generally cannot be extended byte by byte, so each addition becomes another file — and every later read then pays a fixed cost per file.

open as a page

A team choosing one file shape for the datasets it hands between its own steps keeps arguing file size: what should decide it instead?

level: principalimportance: should knowfreq 38%

basics

~20 s

File size is the difference that decides least. Decide on who must be able to open the file without your code, how much of it a typical read touches, whether types must be guaranteed rather than re-established, and how the dataset is produced.

open as a page