A column held as integer codes with a declared list of permitted values is written to text — what comes back?
answer
- two pieces travel, only one is a value
- the list is a declaration, not a description
- text has rows, not declarations
- rebuilt lists narrow to what appeared
basics
~20 sThe values come back; the declaration does not. Text writes out the values the codes stood for, so the reader sees ordinary characters, and any list rebuilt afterwards holds only what this file happened to contain.
solid answer
~50 sA column stored as codes is two things: small integers per row, and one declared list of the distinct values those integers stand for. Re-parsed delimited text — one record per line, every value stored as characters — has a slot for the first and no slot for the second, so the writer spells out the value on each row and the list is simply not in the file. On the read the column is characters like any other, and if anything rebuilds a list from it, that list contains the values that appeared here, in whatever arrangement the tool chooses. A value that was permitted but absent from this extract is gone, and so is any order the declaration carried. A self-describing typed file can record the list; whether its reader hands the column back code-backed or as plain text differs between tools.
go deeper
Know that such a column carries a list of permitted values as well as the values themselves, and that a plain text file has room for only the second of the two.
Explain the narrowing: any list rebuilt on the way back contains what this extract held, so a permitted value that did not appear is gone and two extracts can disagree about the column.
Decide whether the list is load-bearing for a consumer, and if it is, carry it deliberately — in a shape that records it, or as an artefact applied on the read — rather than hoping the extract was complete.
The standing judgement is whether a permitted-value list belongs in the data handoff at all or in a definition the consumers share, and what a team pays when every extract can silently redefine a column.
## What a code-backed column actually holds A **stored-as-codes column** is two pieces of information travelling together: one small integer per row, and one **declared list** of the distinct values those integers stand for. The per-row integers are the data. The list is a declaration — a statement of what this column is allowed to contain, which is not the same as a description of what it currently contains. In some tools the list also carries an order, which is why such a column can sort into a meaningful arrangement rather than an alphabetical one. This leaf is not about why you would store a column that way. It is about which of those two pieces survives a write and a read. ## What the writer can put in a text file **Re-parsed delimited text** is one record per line, every value stored as characters, and every read deciding afresh what each field means. Its slots are fields on lines. A per-row code has a slot; the writer simply prints the value the code stood for, which is the readable and almost always correct thing to do. The declared list has no slot at all. There is nowhere in the file to say *"this column may hold these seven values, in this order"*. | what travelled in memory | what reaches a text file | what a read can restore | |---|---|---| | one integer code per row | the value the code stood for, as characters | the value, as characters | | the declared list of permitted values | nothing | at best, the values present in this file | | the order the declaration carried | nothing | whatever rule the tool applies by default | ## The narrowing, and why it is quiet Suppose the declaration permitted seven values and this extract contains five of them. On the way back, nothing knows about the other two. If a later step rebuilds a code-backed column from the data, it builds a list of five. Two consequences follow, and neither announces itself: - **A permitted-but-absent value becomes indistinguishable from a forbidden one.** The declaration was the only place that distinction lived. - **The set is now a function of the extract rather than of the column.** Two files written on two days from the same source can come back with two different lists, so the same downstream code sees a different column shape depending on which day's extract it was handed. Both are silent. The values in the rows are all correct, the column looks healthy, and the thing that changed is a declaration nobody printed. ## What the order loses Where the declaration carried an arrangement — a severity scale, a size ladder, a set of stages — that arrangement was part of the declaration and not part of any row. Text writes rows. So a rebuilt list falls back to whatever default the tool uses, which is usually the order of first appearance or an alphabetical order, and a report that used to read from lowest to highest now reads in neither. If the order is meaningful, it has to be written as data of its own or restated on the read; it cannot be recovered from the values. ## What a declared shape can and cannot promise A **self-describing binary columnar file** — values stored column by column in their declared types, with the declaration written into the file itself — does have room for this. The column can be recorded as codes plus its list, and a reader can restore it exactly. But two things vary here and both are worth establishing before you rely on either: 1. **Whether the reader restores it as a code-backed column at all.** Some hand the column back as codes with its list; some hand it back as plain values and leave the reconstruction to you; some offer it as an option. The file carrying the information does not oblige the reader to use it. 2. **Whether the tool has such a column type in the first place.** Not every ecosystem models a fixed set of permitted values as a column type, and in one that does not, the question does not arise and there is nothing to lose. ## Handling it in practice If the list is load-bearing — because a consumer validates against it, or because reports must show a category with no rows this week — then either write to a shape that carries the declaration, or make the list an artefact in its own right that travels beside the data and is applied on the read. What you must not do is assume the list survived because the values did. Write the table, read it straight back, and look at the column's list on both sides: the answer takes a minute and it is the same minute whichever tool you are using.
- If all the values came back, why does the missing declaration matter?Because the declaration is a promise about what the column may hold, not a report of what it does hold. Code written against a fixed set now sees whatever this extract contained, and a value that was permitted but absent this week cannot be told apart from one that was never allowed.
- Does a self-describing typed file guarantee the column returns code-backed?No. The shape can record that the column was codes plus a declared list, but whether a reader rebuilds it that way, hands you plain values, or offers it as an option differs between tools. The file can carry the information; using it is the reader's decision, so establish which yours does.
- What happens to an order the declaration carried?Ordinarily it is lost. Text writes rows, and the arrangement was part of the declaration rather than part of any row, so a rebuilt list falls back to the tool's default rule. If the order is meaningful, it has to travel as data of its own or be restated on the read.
saying these in an interview costs you the question
- Says the column round-trips fine because every value came back.
- Treats a rebuilt list as equal to the declared one.
- Assumes a declared order survives because the values do.
- Thinks a self-describing file always hands the column back code-backed.
- Believes every tool models a fixed set of values as a column type.