Why is adding a field to a million-row table usually cheap while appending rows is not, and how much does that vary?
answer
- the holding is organised by field
- a field is one run; a record crosses all of them
- direction is safe, degree is not
- a shared slab can reverse the degree
basics
~20 sBecause the holding is organised by field: a new field is one more run and reads none of the others, while a new record has to place a value into every field's storage. That direction is stable; the degree ranges from one extra chunk to a full rebuild.
solid answer
~40 sA column-major table keeps fields, not records, so the two operations are not mirror images. Adding a field is one more run of values and the existing fields are untouched. Appending records means reaching into **every** field's storage, because that is where the pieces of a record have to go. The direction holds in every design in this family; the degree does not. A chunked holding absorbs a batch of records as one more chunk per field and copies nothing that already exists, which makes the append nearly free. A holding that consolidates same-representation fields into one slab may rebuild that whole slab for a single added field, so there the *field* add is the expensive one. The rule that survives everywhere is: do it once in a batch, not once per record.
go deeper
Remember the direction and its reason: the table keeps fields, so a new field is one more run while a new record has to be written into every field. Batch your appends.
Explain both halves — the direction from column-major storage, and the degree from whether the field is one run, a sequence of chunks, or a slice of a shared slab.
Recognise the pattern in real code: a per-record append loop, or twenty successive field adds, and know the repair is to accumulate cheaply and attach once rather than to tune the loop.
Worth a house rule: build tables at boundaries, not incrementally in business logic. It removes a whole class of surprise whose cost depends on a storage design nobody chose deliberately.
## Where the asymmetry comes from In a **labelled table** — a rectangle whose columns are named and typed one at a time — the values are laid down field by field. One field's values are kept together; one record's are not. Every cost in this question falls out of that one arrangement. - **Adding a field** introduces a new run of values alongside the existing ones. The other fields are not read and, in most holdings, not touched. - **Appending a record** has to put one value into each field's storage, because that is where each piece of the record belongs. A ten-field table means ten places to write; a four-hundred-field table means four hundred. So the *direction* is a property of column-major storage itself, not of any particular tool: fields are the cheap axis to extend, records the expensive one. ## What a field add actually does | holding | what an added field costs | |---|---| | one run per field | one new allocation; nothing existing is read | | chunked field | one new sequence of runs; nothing existing is read | | slab per representation | if the slab already holds that representation, it may be allocated bigger and every existing field of that form copied in | That last row is why "adding a field is cheap" must not be said flat. In a consolidating design a narrow derived field can move an amount of memory proportional to the whole table. ## What a record append actually does | holding | what an appended batch costs | |---|---| | one run per field | every run has to make room; where it cannot, a bigger allocation and a copy of what was already there | | chunked field | one more chunk per field; nothing existing is copied, so the batch is close to free | | slab per representation | the slab has to grow, which generally means a new allocation and a copy | And a design that is immutable by default produces a new holding either way, sharing the runs it did not touch, so the absolute costs are lower than a copying design in both directions while the same asymmetry still shows through in which runs had to be rebuilt. ## The degree, which is what people actually argue about Put the two tables together and the honest statement has two halves: 1. **The direction is safe.** In any column-major holding, extending by a field reaches into fewer places than extending by a record. 2. **The degree is not.** A record append ranges from one extra chunk to rebuilding every field's storage, and a field add ranges from one allocation to rebuilding a shared slab. Which one you get is decided by the storage design, which the logical picture does not show you. Someone who says only the first half sounds right and will still be surprised in production. Someone who says only the second half cannot explain why anyone believes the first. ## The rule that survives every design Appending **one record at a time in a loop** is bad under all three holdings, for three different reasons: - under one run per field, each append may trigger room-making across every field; - under a chunked field, each append becomes its own tiny chunk and the fragmentation is charged to everything downstream; - under a shared slab, each append can force the slab to be reallocated. The repair is the same in each case and does not require knowing which holding you are on: **accumulate the records somewhere cheap and attach them in one batch**. A plain list of records, or a set of per-field sequences, costs almost nothing to grow; converting once at the end pays the table-building cost exactly once. The same logic applies to fields: if twenty derived fields are coming, compute them all and attach them together rather than twenty times. ## How to answer this in an interview Give the mechanism first, then the qualification, then the rule: - *mechanism* — the holding keeps fields, so a field is one more run and a record touches every run; - *qualification* — the degree depends on whether a field is one run, a sequence of chunks, or a slice of a shared slab, and in the last of those a field add can be the expensive one; - *rule* — batch in both directions, because every holding punishes the per-item loop. That structure shows you know the model, know its edges, and know what to do with it — which is more than a candidate who only knows that "appending rows to a table is slow".
- Is there a holding where appending a batch of records is genuinely cheap?Yes. Where a field is an ordered sequence of separately allocated chunks, the batch becomes one more chunk per field and nothing that already exists is copied. The cost is deferred rather than removed: the field is now more fragmented, so later sweeps cross more boundaries and any consumer wanting one contiguous buffer pays for the concatenation.
- Can adding a field ever cost more than appending a batch of records?In a holding that packs every same-representation field into one slab, yes. Adding another field of a form the slab already holds can mean allocating a bigger slab and copying the existing fields into it — work proportional to the whole table for one narrow field. That is the case the usual advice gets wrong.
- Why is a loop that appends one record at a time bad in every holding?Each holding punishes it differently — repeated room-making, one tiny chunk per record, or a slab reallocation per record — but all three charge a per-record cost that scales with the table already built. Accumulating cheaply and attaching once converts that into a single table-building step.
saying these in an interview costs you the question
- Says appending records is slow because records are stored as objects
- Claims every design pays the same price for a batch append
- Treats "adding a field is cheap" as true regardless of how fields are stored
- Appends one record per iteration and expects a batch's cost
- Assumes a field add can never touch another field's storage