A team choosing one file shape for the datasets it hands between its own steps keeps arguing file size: what should decide it instead?
answer
- size is the cheapest line
- count the readers, not the bytes
- how much does a read touch?
- one shape opens anywhere
- you can pay for both, with an owner
basics
~20 sFile size is the difference that decides least. Decide on who must be able to open the file without your code, how much of it a typical read touches, whether types must be guaranteed rather than re-established, and how the dataset is produced.
solid answer
~60 sStorage is the cheapest line in this comparison, and it is a one-off; everything else is paid on every read, by everyone, forever. The four questions that actually decide are: **who opens it** — one shape opens in anything at all, including a plain viewer and a person mid-incident, the other needs something that understands it; **how much of it a read touches** — a shape that stores columns apart earns its keep on wide datasets read narrowly, and earns nothing where every read is the whole file; **whether types must be guaranteed** or may be established afresh by each consumer's code; and **how it is produced** — records arriving continuously push toward the shape that concatenates, or toward a file count and a consolidation rule. It is also worth saying out loud that either shape can be regenerated from the other, so this is a cheaper decision to revisit than the argument suggests — and that writing both is defensible, provided one is named the source of truth and the other is always derived.
go deeper
Know that the two shapes differ in more than size: one opens in anything, the other carries its own types. That pair of properties is the start of every version of this argument.
Be able to list what recurs and what is one-off. Storage is paid once; what a read touches and whether types are guaranteed are paid on every read by every consumer.
Bring evidence rather than preference: what the read pattern actually is, who reads the dataset today, and what it would cost to change the shape once those readers exist.
Own the decision as an interface commitment. Name the default and its reasons, decide whether both shapes are worth writing, and if so appoint the source of truth before anyone needs it.
## Why size is the weakest argument Bytes on disk is the easiest property to measure and the least consequential of the four, which is exactly why teams argue about it. Three things undercut it: - **It is paid once.** A dataset is written once and read many times; a one-off difference in stored bytes competes against a difference paid on every read. - **It is data-dependent.** Which shape is smaller depends on what the values look like, so neither side of the argument can honestly claim a general winner. - **It predicts nothing else.** Bytes on disk do not tell you what a read touches, what the file guarantees, or who can open it — and those are the properties you will live with. The move worth learning here is refusing the framing: not "which is smaller" but "which of these differences will this team pay for repeatedly". ## The four properties that actually decide | Question | Favours re-parsed delimited text | Favours a self-describing typed shape | |---|---|---| | Who must open it without your code? | unknown, heterogeneous, or a person | your own steps, all using one toolkit | | How much does a read touch? | the whole dataset, every time | a few columns of many, often | | Must types be guaranteed? | no; each consumer may decide | yes; the guarantee must travel with the bytes | | How is it produced? | records arrive continuously, appended | written in batches as finished units | None of these is a tiebreaker on its own. The first is the one teams undervalue, because the readers they are imagining are the ones they have written. The readers that matter are the ones they have not: an analyst with a different toolkit, a partner team, a person at two in the morning trying to see what is in the file that broke the run. ## The commitment is about consumers, not files A file shape looks like a property of a file and behaves like a property of an interface. Once a dataset is published in a shape, readers are written against it, and each one is a small vote against ever changing it. That is what makes this a standing decision rather than a per-file choice, and it has two consequences worth stating in the design discussion: 1. **The cost of the shape grows with adoption**, not with the data. A shape chosen badly for two consumers is cheap; the same shape with twenty consumers is a project. 2. **Interoperability is easy to give up and expensive to buy back**, because buying it back means either writing a converter everyone must find, or changing every consumer. ## Paying for both Writing both shapes is a legitimate answer, and it is the usual one at a boundary: the typed shape for the pipeline's internal steps, where the readers are known and the reads are narrow, and a text copy at the edge, where they are not. The cost is not the storage. It is that two artefacts can disagree, and that nobody decided which is right. Make it safe with two rules: - **One is the source of truth**, named in writing, and the other is always derived from it by the same job that produced it. - **Nothing edits the derived copy**, ever — if it is wrong, the fix goes into the source and the copy is regenerated. Without both rules you have not bought interoperability; you have bought an argument about which number is correct. ## How reversible is it really More reversible than the argument implies, and asymmetrically so. Either shape can be regenerated from the other, so no data is trapped. What is not reversible is the set of consumers: converting the files is a job, while converting everyone who reads them is a negotiation. So treat the decision as cheap while the dataset is internal and few things read it, and as expensive the moment it is published to anyone who is not in the room. ## A defensible default Where the team genuinely has no signal yet, a common and defensible default is: the typed shape for anything handed between the team's own steps, and text at the boundaries where the readers are unknown or human. It gets the guarantee where the guarantee is worth having, and the openness where openness is the whole point. But state it as a default with reasons attached, not as a house style — because the moment one of the four questions above has a clear answer for a particular dataset, that answer should beat the default, and a team that cannot articulate why it chose its shape will not notice when the reason expires.
- When is the text shape the right commitment even inside a pipeline?When the readers are unknown or use different toolkits, when a person must be able to open the file during an incident, when every read consumes the whole dataset anyway, and when nothing downstream depends on a type being guaranteed rather than re-established.
- What is the standing cost of writing both shapes?Two artefacts that can disagree, twice the write, and a question nobody answers in advance: which one is right. It is defensible with one named source of truth, the other always derived by the same job, and nothing ever editing the derived copy.
- How would you make this decision when the read pattern does not exist yet?Pick a default with reasons written down, keep the dataset internal while few things read it, and revisit once real read patterns appear. The data is convertible either way; the consumers are what becomes expensive, so delay publishing more than delaying the choice.
saying these in an interview costs you the question
- Decides on bytes on disk and stops there.
- Commits to a typed shape without asking who else opens the files.
- Treats readable-by-anything as worthless because the team writes every reader.
- Argues the choice is irreversible when either can be regenerated.
- Writes both shapes without naming which is the source of truth.