skip to content

A process must keep adding records to a dataset on disk: what does appending cost on each of the two file shapes, and what accumulates?

level: seniorimportance: should knowfreq 46%

answer

  1. one shape concatenates, one finalises
  2. additions become files, not bytes
  3. fixed setup per file, every read
  4. rewrite many small into few large

basics

~20 s

Delimited text takes new records by concatenation, so appending is nearly free. A finalised typed file generally cannot be extended byte by byte, so each addition becomes another file — and every later read then pays a fixed cost per file.

solid answer

~60 s

**Re-parsed delimited text** — one record per line, values written as characters — is a stream of lines, so adding records means writing more lines on the end. Nothing has to be rewritten, which is exactly why it is cheap. **A self-describing binary columnar file** — values stored column by column in their declared types, with the declaration written into the file — is completed as a unit: raw bytes tacked on the end are not part of what the file declares itself to contain, so a reader will not see them. Designs vary in what they offer instead: some expose a write that adds a further segment to a file, and it is common to treat a directory of files as one dataset and simply add another. What accumulates in that second case is **files**, and a read of a dataset made of forty thousand tiny files pays a fixed setup per file — open it, read its declaration, plan the read — that does not shrink because the file is small.

go deeper

for a junior

Know that adding records to a text file is writing more lines on the end, and that a typed binary file is written as a finished whole, so growing it is not the same easy operation.

for a middle

Explain why the typed file resists byte-level appending, and name what teams do instead: add another file, add a segment, or rewrite. Then say what a growing file count does to reads.

for a senior

Diagnose the small-file pattern from the symptom — reads slowing while the data barely grows — and propose batching plus a triggered consolidation, with an owner, rather than a one-off cleanup.

for a principal

Decide where the cost should sit. The writer creates it and readers pay it, so the consolidation rule and its owner are a standing commitment, not an operational chore to be discovered twice a year.

## Two shapes, two meanings of "adding records" Appending records to a file on disk sounds like one operation, and it is two. On **re-parsed delimited text** it is concatenation: a record is a line, a file is a sequence of lines, so more records means more lines on the end and nothing already written has to change. On **a self-describing binary columnar file** it is not concatenation at all, because the file is not a sequence of records — it is a completed object whose own declaration says what it contains and how it is laid out. Bytes appended outside that declaration are, from a reader's point of view, not in the file. This is one of the few places where the two shapes differ in **what operations exist**, rather than in how much they cost, and it is the property most often discovered late. ## What designs actually offer instead It would be wrong to say a typed dataset cannot grow. What varies is the mechanism: - **Add another file.** The commonest answer: readers routinely treat a set of files as one dataset, reading each and stacking the results. Growth becomes a matter of file count. - **Add a segment to the file.** Some writers expose an append that extends the file properly, rewriting or extending its declaration so the new records are genuinely part of it. - **Rewrite the whole file.** Always available, always proportional to the data already there, and the reason nobody does it per record. What is *not* generally available is the cheap one: opening the file and writing bytes on the end. ## The cost lands on the reader, not the writer The writer that creates one small file per addition is having a good time. The cost shows up later and elsewhere: | | One large file | Forty thousand small files | |---|---|---| | Cost to add records | rewrite or extend | trivial: write a new file | | Fixed setup per read | paid once | paid forty thousand times | | Declarations to reconcile | one | one per file, and they must agree | | Effect of a per-file saving | applies to real data | applies to almost nothing | The fixed setup is the heart of it: opening a file and reading its declaration before any value is produced is a cost per **file**, not per **record**. A dataset whose file count grows with time therefore gets slower to read at a rate that has nothing to do with how much data it holds. And because the files must agree about their columns for a reader to stack them into one table, a producer that quietly changes what it writes turns a cheap append into a read that fails or a column that arrives wider than anyone expected. ## The remedy, and who owns it The standard answer is to periodically rewrite many small files into fewer large ones, on a rule rather than by hand: 1. **Pick the trigger** — a file count or a total size, not somebody noticing the reads got slow. 2. **Name an owner.** The cost is paid by readers and created by a writer, so it belongs to neither by default and is therefore nobody's until it is assigned. 3. **Make the rewrite safe to interrupt**, since it is the one operation that touches data already published. ## The other side of the cheap append Text's easy append is not free of consequences, and a senior answer says so: - **The naming line.** Concatenating two complete text files leaves the second file's column-naming line sitting in the middle of the records, where the next read treats it as data. - **No agreement is enforced.** Because the shape declares nothing, a producer that starts writing an extra field, or the same fields in a different order, appends perfectly happily and the damage is discovered downstream. - **Partial writes look plausible.** An interrupted append leaves a file that still looks like lines of text, whereas an interrupted typed write leaves something a reader will refuse outright — which is unpleasant, but at least it is loud. ## How to decide If records arrive continuously and the dataset is read whole, the text shape's append is genuinely the cheaper design and the small-file problem never appears. If records arrive continuously and the dataset is read selectively and often, you are choosing between a growing file count and a periodic rewrite, and the right move is usually to accept both — batch the additions so that each file is worth opening, and schedule the consolidation instead of hoping the problem stays small. What you should never do is pick the shape on how convenient the write is, because the write happens once per batch and the read happens forever.

  • How is a directory of many files read as one dataset, and what does that require of them?
    A reader opens each file and stacks the results into one table. That requires the files to agree about their columns, and it costs one fixed setup per file — opening it and reading its declaration — before any value is produced.
  • If every addition creates a file, what keeps the read cost bounded?
    Batching additions so each file is worth opening, plus a consolidation that rewrites many small files into fewer large ones on a size or count trigger. Both need an owner, since the cost is created by the writer and paid by readers.

saying these in an interview costs you the question

  • Says a finished typed file can be extended by appending bytes.
  • Treats ten thousand small files as equivalent to one large one.
  • Concatenates text files without accounting for the naming line.
  • Believes no design lets a typed dataset grow at all.
  • Picks the shape on write convenience while every read pays.