skip to content

After two write-and-read-back cycles through text, a table carries two extra leading columns of numbers — why?

level: middleimportance: should knowfreq 55%

answer

  1. something was written that was not a column
  2. flat text has no slot for row identity
  3. the reader sees data, not an identity
  4. one extra column per round trip

basics

~10 s

Each trip wrote the table's row labels out as an ordinary leading column, and each read brought them back as plain data. Two trips, two columns. Tables with no row-identity concept never show it.

solid answer

~50 s

Some tables carry a row identity alongside their columns — a label per row that is not itself one of the columns. Flat text has no slot for it, so a writer that does not want to drop it emits it as an ordinary leading column, often with no name. On the way back the reader has no way to know that column was ever special: it becomes data, and the table receives a fresh set of positional labels behind it. Do it again and you get a second such column. Three things close the loop: do not write the labels when they are only positions, or name that column and tell the reader to promote it back, or write to a shape that keeps per-column metadata. Tools whose tables have no row identity never see this, because nothing was carried.

go deeper

for a junior

Recognise the symptom: an extra leading column of numbers nobody created means something that was not a column got written as one. Do not just delete it and move on.

for a middle

Explain the mechanism both ways — what the writer had nowhere to put, and why the reader cannot tell a written-out identity from a genuine unnamed first field — and give the fix on whichever side is cheaper.

for a senior

Own the handoff: decide whether those labels carry information at all, make the choice explicit on both sides, and add the write-then-read-back comparison that catches a column-per-generation drift before a consumer does.

for a principal

The standing question is whether row identity should be part of a file handoff at all, or whether every exchanged file should carry its key as a named column so that no two tools have to agree about a concept only some of them have.

## Something was written that was not a column In some tools a table carries a **row identity** in addition to its columns: one label per row, stored apart from the data and not addressed as a field. In other tools a table is columns and nothing else, and a row is found only by its position. That difference is the whole reason this symptom exists for some people and is incomprehensible to others. When such a table is handed to **the writer** — the call that turns a table back into bytes — and the target is **re-parsed delimited text**, one record per line with every value stored as characters, the writer faces a shape with exactly as many slots as there are fields. There is no per-row slot beside them. It has two options, and both are defensible: - **drop the labels**, which is lossless when they were only the positions the rows happened to occupy, and destructive when they were an identifier; or - **emit them as an ordinary leading field**, which preserves the values at the cost of turning them into data. The second is the common default, and it is where the extra column comes from. ## Why the reader cannot tell On the way back, **the reader** — the call that turns a file into a table in one step — sees a line of fields. Nothing in the file says that the first of them used to be a row identity rather than a measurement. If the writer emitted no name for it, the reader sees an unnamed first field, which is also exactly what a legitimately unnamed first column looks like. Guessing would be wrong for somebody, so readers leave the decision to you: the field becomes a column, and the table is given a fresh identity, usually the row positions. The result is a table that is **wider by one** and whose row identity is no longer the one you had. ## Why it compounds The new identity is real, so the next write faces the same choice and makes the same decision: | trip | what the file gains | what the table comes back as | |---|---|---| | first | the original labels, as a leading field | original columns plus one column of old labels, with fresh positional labels | | second | the fresh positional labels, as another leading field | original columns plus two columns of old labels, with fresh positional labels again | One per round trip, forever, and nothing anywhere raises an error. Somebody eventually deletes the leading columns by hand, in one of the jobs that consumes the file, and the problem survives in the others. ## Three ways to close the loop 1. **Do not write them.** If the row labels are only positions, they carry no information and writing them buys a column on every trip. Tell the writer to leave them out, and the round trip is stable. 2. **Name them and promote them back.** If the labels do carry something — an identifier, a key, a moment — then write them under a real column name, and on the read tell the reader to promote that named column back to the table's row identity. Both halves are required: naming alone gives you a well-named data column, and promoting a column nobody named is guesswork. 3. **Write a shape that keeps per-column metadata.** A self-describing typed file — values stored column by column in their declared types, with the declaration written into the file itself — has somewhere to record which column carried the identity, so a round trip through it can return what went in. ## Where this problem does not exist at all In tools whose tables have no row-identity concept, there is nothing beside the columns to write, nothing to lose and nothing to accumulate. A learner who has only used those tools will find the symptom baffling, and a learner who has only used the others will assume every tool has it. Both are wrong, and the useful statement is conditional: *a table that carries row identity, written to a flat text shape, either loses that identity or gains a column.* ## Noticing it early The cheapest detection is the round trip itself: write the table, read it straight back in the same process, and compare the column names on both sides. A table that gained a column has told you everything you need to know, in one step, before any consumer built a pipeline on top of a file that grows a column per generation.

  • Would writing the labels under a name instead of leaving the column unnamed be enough?
    Only half of it. A name makes the column legible and stops the pile-up looking mysterious, but the reader still hands it back as an ordinary column unless it is told to promote that column to the table's row identity. Both sides have to agree, or you have simply renamed the problem.
  • Why does the reader not just treat a leading unnamed column as the labels?
    Because it cannot distinguish labels that were written out from a genuine first field that happens to have no header. Guessing would corrupt every file with a legitimately unnamed first column, so readers leave the promotion to you and take the safe reading.
  • When are the labels worth writing at all?
    When they carry something the columns do not — an identifier, a moment, a key the consumer needs to join on. When they are only the positions the rows happened to occupy, writing them costs a column on every trip and tells the consumer nothing it could not recompute.

Printing a spreadsheet with its row numbers shown down the side, then typing the printout back in as a fresh sheet. The old row numbers are now a column of data, and the new sheet has its own row numbers down the side. Print and retype that, and you have two columns of old numbers.

saying these in an interview costs you the question

  • Calls the extra column a bug in the reader.
  • Deletes the leading column after each read instead of fixing the write.
  • Assumes every tool's tables carry row labels to lose.
  • Says the labels are preserved because the values are still there.
  • Thinks naming the column is enough without promoting it back on read.