A column of 5 million text values is measured two ways and the figures differ tenfold — what storage form makes that possible?
answer
- two storage forms, not one
- the buffer may hold only references
- shallow count stops at the slots
- following references adds per-value overhead
- contiguous bytes plus a position per row
basics
~20 sOne reference per row, with the values allocated separately elsewhere. The column's own buffer is a fixed slot per row, so a shallow count stops there; following the references adds every value's bytes plus its per-allocation bookkeeping.
solid answer
~50 sText has two storage forms in wide use. One keeps a fixed-width reference per row pointing at a value allocated somewhere else; the other keeps all the characters in a single contiguous block of bytes with a position recorded per row. Under the first, the column's own buffer is just the slots, so a shallow measurement reports a slot size times the row count and never moves when the text gets longer. The **deep footprint figure** — the measurement that follows the references and counts what they point at — adds every value's characters plus whatever each separate allocation costs in bookkeeping, and for short values that bookkeeping is often larger than the text itself. Under the contiguous form the two measurements nearly agree, because there is nothing to follow. So before quoting a byte figure for a column of text, say which storage form it uses and which measurement produced the number.
go deeper
Recall that a column of text may hold only a reference per row, with the value itself somewhere else. That is why a byte figure for such a column can be far smaller than the truth.
Explain both storage forms and what each measurement reports, and do the arithmetic: slots times rows against characters plus bookkeeping times rows. Say which measurement produced any number you quote.
In a sizing exercise, flag the text columns as the terms you cannot bound without following references, and say so before anyone plans around your total. The failure mode is a confidently round figure nobody questioned.
The choice between the two forms is a standing decision about every table your teams build. One contiguous buffer per column trades a slightly rigid write path for a footprint that is predictable from the data, which is what makes capacity arguments possible at all.
## Two ways a column of text is stored A column's **representation** — the one physical form every value in it is stored in — takes two shapes in wide use for text, and they cost completely different things. - **One reference per row, values allocated separately.** The column's own buffer is a fixed-width slot per row. Each slot holds a reference to a value sitting somewhere else, and each of those values carries its own bookkeeping: a length, a small header, and whatever the allocator rounds up to. - **One contiguous block of bytes, with a position recorded per row.** Every character of every value sits end to end in one buffer, and a second small buffer records where each row's value begins. There is one allocation for the whole column rather than one per value. Neither of these is *the* way text is stored. Which one you have is a property of the design in front of you, and several tools offer both for the same column. ## What each measurement reports - A **shallow** count reports what the column's own buffer occupies — the slots, and nothing further. - The **deep footprint figure** is the measurement that follows the references and counts what they point at, plus the per-allocation bookkeeping. Under the contiguous-buffer form these two figures are nearly the same, because there is nothing to follow. Under the reference form they are different numbers answering different questions, and the shallow one is smaller by a large factor. ## The arithmetic, worked Five million text values averaging thirty characters each. Take an illustrative fifty bytes of per-allocation bookkeeping — the exact figure is a property of the runtime, not something to memorise. | Storage form | What is counted | Bytes | |---|---|---| | References, shallow | 5,000,000 slots × 8 | 40,000,000 | | References, deep | slots + (30 characters + 50 bookkeeping) × 5,000,000 | ~440,000,000 | | Contiguous buffer | 30 characters × 5,000,000 + one 4-byte position per row | ~170,000,000 | Forty megabytes against four hundred and forty is the tenfold gap. It is not a rounding error and it is not conservatism; the two numbers are answers to two different questions. ## Why the gap is a factor rather than a margin Notice where the bytes go under the reference form. For short values **the bookkeeping is larger than the text**. A thirty-character value paying fifty bytes of overhead spends more on being a separate allocation than on its own content — and there are five million of them. That is why the reference form's real cost scales with the row count almost as steeply as with the total volume of text. The contiguous form removes exactly that: one allocation instead of five million, and a position per row instead of a reference plus a per-value header. Its estimate — the bytes of the characters, plus a small fixed cost per row — is honest, which is the whole reason the form exists. ## What this means for an estimate 1. Before quoting a figure for a column of text, say which of the two storage forms it uses. If you cannot say, you cannot quote the figure. 2. If it is the reference form, say which measurement produced your number. A shallow count of such a column is not wrong; it is answering a question nobody asked. 3. A shallow figure under the reference form is recognisable on sight: it is exactly a fixed slot size times the row count, and it does not move when the text gets longer. If your "column size" is suspiciously round and suspiciously stable, that is what you are looking at. ## Where this leaves the per-column figure The arithmetic that works so cleanly on a fixed-width numeric column — declared width times rows — is not *wrong* for a column of text under the reference form. It is a measurement of the slots. Treat it as a floor. The honest per-column figure is the slots plus the deep part, and it is the only per-column figure in a table that you cannot obtain without following references. That asymmetry is worth stating when you hand the number to someone: one column of text can be larger than every numeric column in the table put together, and nothing in the shallow arithmetic hints at it. ## The claim to keep bounded "A column of text costs roughly the length of the text, per row" is true of the contiguous-buffer form and false of the reference form, whose own buffer does not know how long anything is. Stated flatly it describes one design and misleads everyone using the other. The usable version names both: under a contiguous buffer the cost really is the characters plus a small per-row position; under one reference per row the cost is a slot per row in the column *plus* an entire separately allocated value, with its own overhead, for every single row.
- Without naming a tool, how would you tell which of the two storage forms a column of text is using?Compare the shallow figure against the figure that follows references. If the shallow one is exactly a fixed slot size times the row count — and stays there when the values get longer — you have one reference per row. If the two measurements agree closely and the figure tracks the total volume of characters, you have one contiguous buffer plus positions.
- Does the contiguous-buffer form make the per-column estimate exact?Close to it. The cost is the bytes of the characters plus one small position per row plus a fixed per-column header, so the error is bounded and small, and the figure moves with the data as you would expect. Under the reference form the error is unbounded from the shallow side, because nothing in the slots records how much is on the other end.
- Does a column of very long text values behave differently under the two forms?Yes, the gap narrows. The per-allocation bookkeeping is roughly fixed per value, so as the average value grows from tens of characters to thousands it becomes a small share of the total and the two forms converge. The reference form's penalty is worst for columns of many short values, which is the common case for codes, labels and short names.
A cloakroom keeps one numbered ticket per coat on a rack. Counting the tickets tells you exactly how many coats there are and exactly what the rack weighs — and nothing whatever about the coats. A column that stores one reference per row is the ticket rack; the text is in the cloakroom.
saying these in an interview costs you the question
- Quotes a shallow byte count as the column's real cost.
- Assumes every design stores a column of text the same way.
- Says text costs about its character count per row, full stop.
- Ignores the per-allocation bookkeeping on short text values.
- Believes the gap is a small margin rather than a factor.