skip to content

A table gives each column its own type — in what three ways can the values behind one column actually be held?

level: middleimportance: should knowfreq 48%

answer

  1. the logical picture hides the storage
  2. one run, many chunks, one shared slab
  3. can you point at the field?
  4. adding a field may rebuild a slab

basics

~20 s

Three holdings are common: one contiguous run per field; an ordered sequence of separately allocated chunks; or one slab per representation, in which a field is a slice. The logical picture is identical in all three; the costs are not.

solid answer

~50 s

The logical picture — one named field, one **column representation** — says nothing about where the values sit. Three physical holdings all satisfy it. **One run per field** gives the field a single allocation you can point at. **A chunked column** holds it as an ordered sequence of separately allocated runs, so appending a batch is one more run and nothing existing is copied, but there is no single buffer to hand out. **One slab per representation** packs every field sharing a form into one two-dimensional allocation, so sweeping all of those fields is one pass, while a field is a slice that does not own its memory and adding another field of that form can force the slab to be rebuilt. Any cost claim about "a column" is really a claim about one of these three.

go deeper

for a junior

Know that the type is per field and that the values of one field are kept together. The detail to take away is that "kept together" has more than one implementation.

for a middle

Name the three holdings and give one cost each: a run you can point at, chunks that make batch appends cheap, a shared slab that makes a field add expensive. That trio is the answer.

for a senior

Demonstrate that you attach preconditions to cost claims and that you would time the operation across width and length rather than reasoning from a holding you assumed.

for a principal

The standing call is whether a codebase may depend on any of this. Depending on contiguity or on cheap appends ties the code to one holding, which is a portability cost worth naming before it is paid.

## The logical picture hides the storage In a **labelled table** — a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own — every field declares one **column representation**: the single physical form every value in it is stored in. That sentence is about typing. It does not say whether the field is one allocation, several, or part of a bigger one, and those are three genuinely different designs shipping in this family today. This matters because almost every cost claim people make about a table is really a claim about the storage, dressed up as a claim about the logical column. ## One run per field The field is a single allocation holding its values back to back. - There is a buffer, and you can point at it. Anything that wants the field contiguously already has it. - Adding a field is one new allocation and does not read any existing field. - Appending records has to make room in **every** field's run, which can mean allocating a bigger one and copying what was already there. - The field owns its memory: freeing it frees the values. ## A chunked column The field is held as an ordered sequence of separately allocated runs rather than as one buffer. - Appending a batch of records is one more chunk per field and copies nothing that already exists. This is the case where "appending rows is expensive" is weakest. - There is no single buffer. Handing the field to a consumer that demands contiguity means allocating a new run and copying every chunk into it. - Anything sweeping the field crosses chunk boundaries, so the sweep is a loop over runs rather than one run. - Appending **one record at a time** is the failure mode: you get a long sequence of tiny chunks, and everything downstream pays for the fragmentation. ## One slab per representation Every field sharing a representation is packed into a single two-dimensional allocation, and a field is a slice of it. - Operations touching all the same-representation fields at once are a single pass over one allocation. - A field does not own its memory. The slab stays alive as long as any slice of it does, so dropping one field frees nothing on its own. - Adding another field of a representation the slab already holds may mean allocating a bigger slab and copying the existing fields into it — the one case where a field add is the *expensive* direction. - Whether the slice itself is contiguous depends on the slab's own layout; where the slab runs field by field, the slice is contiguous and can be handed out as a window. ## Side by side | | one run per field | chunked field | slab per representation | |---|---|---|---| | is the field one allocation? | yes, and it owns it | no, a sequence of them | it is a slice of a shared one | | append a batch of records | make room in every run | one more chunk per field | extend or rebuild the slab | | add a new field | one allocation | one allocation | may rebuild the whole slab | | hand out a contiguous buffer | already available | allocate and copy | available if the slab runs field by field | | drop one field | frees its values | frees its chunks | frees nothing until the slab goes | ## Why you should not assume The sentence *"a table stores each column as one contiguous run of values"* is the single most commonly asserted thing in this subject and it is true of exactly one of the three. Two habits follow from that: 1. **Attach the precondition to the claim.** "Appending a batch is cheap" is a statement about a chunked holding. "Adding a field is one allocation" is a statement about the two non-consolidating holdings. Said flat, each one is wrong somewhere. 2. **Measure rather than reason, when the cost matters.** The holdings are not visible from the logical picture, so the way to find out which one you are on is to time the operation you care about as the table gets wider and as it gets longer, and see which dimension the cost tracks. ## What this does not decide The storage decides physical costs, not semantics. The same table, the same field names, the same values and the same per-field forms are on offer in all three; what changes is which operations are an allocation, which are a copy, and which are free. Keep the two apart in your own explanations, because an interviewer asking "what does that cost?" is asking about the storage, and an interviewer asking "what will it return?" is not.

  • Why does it matter whether a field is one allocation or a sequence of chunks?
    Three things change. Appending a batch is one more chunk instead of making room in an existing run. Handing the field to something that demands contiguity now costs an allocation and a full copy. And a sweep over the field is a loop over runs rather than one run, which matters when the chunks are small.
  • What does it change if a field is a slice of a slab shared with other fields?
    Ownership and lifetime. The slab is freed only when its last slice is, so dropping one field reclaims nothing. Adding another field of the same form may force a bigger slab and a copy of the existing ones. In exchange, anything sweeping all the same-form fields is one pass over one allocation.

saying these in an interview costs you the question

  • Assumes every table keeps each field as one contiguous run
  • Treats a chunked field as a single buffer that can be handed out as-is
  • Says the storage design is an implementation detail with no cost consequences
  • Believes a field that is a slice of a shared slab can be freed on its own
  • Confuses how the values sit in memory with how a file arranges them