skip to content

What Did Not Survive Being Written

Write a table out, read it back, and it is not the same table. Which of the types, the labels, the absent markers and the last decimal digits survive is decided by what you wrote it to.

on this pageshow

questions

4

A table with declared column types is written out as one-record-per-line text and read back — what is no longer guaranteed?

level: juniorimportance: must knowfreq 70%

answer

  1. the shape decides, not the tool
  2. characters carry no declarations
  3. a column name is not a column type
  4. types, labels, markers, lists, digits

basics

~20 s

The types are not guaranteed. Text stores characters only, so the reader must decide afresh what each field is. A file that carries its own type declarations restores them; flat text has nowhere to put them.

solid answer

~50 s

What survives a round trip is decided by the shape you wrote to, not by the tool that wrote it. Re-parsed delimited text — one record per line, every value stored as characters — carries no declarations at all, so on the way back the reader either decides each column's stored representation from what it sees or leaves everything as characters until you declare it. A self-describing binary columnar file — values stored column by column in their declared types, with the declaration written into the file itself — has somewhere to put that information, and its reader reads it rather than inferring anything. Five things are at stake on every trip: the column types, the table's row labels, the absent markers, any declared list of permitted values, and the last decimal digits. Each survives or does not by shape.

go deeper

for a junior

Be able to say that a plain text file stores characters and nothing else, so the types you had are not in the file and something has to decide them again on the way in.

for a middle

Explain what each shape can physically record — a declaration written into the file against nothing but characters — and name what is at stake besides the types: row labels, absent markers, declared value lists, printed digits.

for a senior

Treat the shape as part of the handoff design. Say which losses you accepted, which you closed by declaring types on the read, and how the next person finds out when one of them changes.

for a principal

The call is what the team standardises on for handoffs between jobs, and what it is worth paying for the legibility of text when every consumer must then restate the types the producer already knew.

## The round trip, and what actually decides it A **round trip** is a write followed by a read: the **writer** — the call that turns a table back into bytes — produces a file, and later the **reader** — the call that turns a file into a table in one step — produces a table from that file. The comfortable assumption is that the second table is the first one. How far from it you land has little to do with which tool you used and almost everything to do with **the shape you wrote to**, because the shape decides what there was room to record. ## Two shapes, very different room **Re-parsed delimited text** is one record per line, every value stored as characters, with every read deciding afresh what each field means. The only things in the file are those characters and, if you asked for one, a first line naming the columns. A name is not a type: that line can say a column is called `amount`; it cannot say the column is a number. **A self-describing binary columnar file** stores values column by column in their declared types and writes the declaration into the file itself. Its reader does not infer anything — it reads the types that are there. | in the table you wrote | after flat text | after a self-describing typed file | |---|---|---| | the column's stored representation | decided again on the read | read from the file's own declaration | | the table's row labels | dropped, or emitted as an ordinary column | carried, where the shape keeps per-column metadata | | absent markers | spelled as some literal token | kept as absence, distinct from any value | | a declared list of permitted values | gone; only what appeared remains | carried where shape and reader both support it | | the last decimal digits | only as many as the writer printed | the stored value, bit for bit | ## "The reader infers the types" is only half true It depends on what the reader had to work with. Over a shape that declares its types, the reader reads them and infers nothing at all. Over a shape that does not, two different behaviours are both common: decide each column from some sample of rows, or leave every column as characters until you declare otherwise. So *"it will read my numbers back as numbers"* is a property of one reader over one shape, not a rule of the subject. ## The five things at stake, every trip - **The column's stored representation** — the single physical representation every value in a column is held as. Text does not carry it, so it is re-decided; a declared shape carries it and there is nothing to decide. - **The table's row labels** — in tools whose tables have a row identity at all. Written to flat text they either vanish or come back as an ordinary column of data. Plenty of tools have no row-identity concept, and for those nothing is lost because nothing was carried. - **Absent markers** — whatever the tool puts in a cell that has no value. Text must spell absence as some literal token, and recovering it means the reader maps that token back. How absence is held in memory also differs between designs: some borrow a floating-point sentinel, which forces a narrow integer column to a wider representation; others carry a separate validity bit alongside each value and leave the width alone. - **A declared list of permitted values** — the list attached to a column stored as small integer codes. Text writes the values, not the list, so what comes back is at best the set that happened to appear in this file. - **The last decimal digits** — text holds what the writer printed. Some writers print the fewest digits that read back as the identical value; others print a fixed number, and then the rest is gone for good. ## Turning the loss into a decision 1. **Say what the file is for.** A file a person reads and a file a later job reloads are different jobs, and only the second has to be faithful. 2. **If it must be text, declare the types on the read.** Stating them up front removes the largest hole — the guess — and costs one declaration kept beside the read. 3. **If fidelity matters more than legibility, write a shape that carries declarations.** Then the reader has nothing to decide, and the widths, the declared list and the exact numbers come back. 4. **Read back what you wrote, once, and compare.** A round trip is cheap to exercise and almost never exercised. ## Why this bites harder than it sounds Almost none of it raises an error. A column that quietly arrived as characters still sorts, still compares, still prints — it just orders text instead of numbers. A cell that was absent and came back as whatever text was printed in its place is now an ordinary value. A number short of its final digits still adds up, to something slightly different. The loss is silent by construction, which is why the shape belongs in the design of the handoff rather than in a consumer's bug report three weeks later.

  • If you hand the reader a full set of column types, is the text round trip now lossless?
    No. Declaring the types closes the guessing, which is the largest hole, but the remaining losses are in the file itself: the row labels are gone or sitting in a column, a declared list of permitted values has narrowed to whatever appeared, and numbers carry only the digits the writer printed. A declaration fixes interpretation, not what was never written.
  • Why is the first line naming the columns not a type declaration?
    Because it names columns and nothing else. It tells the reader how many fields there are and what to call them; it says nothing about whether a field is a number, a moment in time, a code or free text. Column names survive a text trip. Column types do not.
  • Which parts of the loss show up as an error, and which as a wrong answer?
    Almost none of it errors. A column arriving as characters sorts and compares differently, a value that was absent may return as ordinary text, and a number may be short a few digits. Every one of those produces a plausible wrong answer rather than a failure, which is why a round trip is worth exercising deliberately.

saying these in an interview costs you the question

  • Says a write followed by a read always returns the same table.
  • Blames the writer for losing types a text shape cannot hold.
  • Assumes a text reader always restores numbers as numbers.
  • Thinks switching tools fixes what the file shape cannot carry.
  • Treats the header line naming columns as a type declaration.
open as a page

A column held as integer codes with a declared list of permitted values is written to text — what comes back?

level: middleimportance: should knowfreq 40%

basics

~20 s

The values come back; the declaration does not. Text writes out the values the codes stood for, so the reader sees ordinary characters, and any list rebuilt afterwards holds only what this file happened to contain.

open as a page

After two write-and-read-back cycles through text, a table carries two extra leading columns of numbers — why?

level: middleimportance: should knowfreq 55%

basics

~10 s

Each trip wrote the table's row labels out as an ordinary leading column, and each read brought them back as plain data. Two trips, two columns. Tables with no row-identity concept never show it.

open as a page

A nightly export writes a table to text and a reload comparison fails on a few decimal values — what happened, and what test catches this class?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The file holds what the writer printed, not the value, so the reload re-parses a printed decimal. Whether that is lossless depends on how many digits the writer prints. Write, read straight back, and compare.

open as a page