A column holds 10 million values each stored in the same declared 8 bytes — how much does it occupy, and what is left out?
answer
- start from bytes per row
- declared width, not value magnitude
- width times row count
- the packed value buffer only
- text breaks the multiplication
basics
~20 sEighty million bytes: the declared width times the row count. That figure is the packed value buffer only — anything stored beside it, such as an absence mask, the held-once values behind codes, or the table's row labels, is extra.
solid answer
~50 sThis is a fixed-width numeric column: every row occupies the same declared number of bytes, and a value outside that range cannot be represented at all. So the per-column figure is one multiplication — 8 bytes times 10,000,000 rows, or 80,000,000 bytes, about 76 MiB. Nothing about the individual values enters it: a row holding `3` and a row holding three billion each cost 8 bytes, because the width belongs to the column's representation and not to the value. The two levers are therefore the declared width and the row count. What the multiplication gives you is the packed value buffer. It does not include a separate absence mask if the design keeps one, the distinct values held once behind a column of codes, or the table's row labels — and it is the wrong tool entirely for a column whose rows hold references to values allocated elsewhere.
code
pseudocode · 14 lines# per-column figure for a fixed-width numeric column
width_bytes = 8 # declared once, for the whole column
row_count = 10000000
value_buffer_bytes = width_bytes * row_count # 80,000,000
# same arithmetic, narrower declared width
narrow_buffer_bytes = 2 * row_count # 20,000,000
# what the two lines above do NOT contain:
# an absence mask beside the values (~row_count / 8 bytes)
# the distinct values behind a coded column
# the table's row labels
# anything a stored reference points atgo deeper
Recall the one multiplication: bytes per row times number of rows. Know that the width is declared for the column, so how big the numbers are does not change the answer.
Explain why the multiplication is exact for a packed buffer — no per-value header, no reference to follow — and name the things it does not count: an absence mask, the values held once behind codes, the row labels.
Be able to size a real table from its column list in your head, and to say which of its columns have a figure you trust and which have only a floor. The honesty about the floor is the part interviewers listen for.
The lever you own is the declared width itself. Narrowing a column halves its bytes and halves the range it can represent; that trade, made once at the table's boundary, is worth more than any later tuning.
## The arithmetic A column's **representation** is the one physical form every value in that column is stored in. When that representation is a **fixed-width numeric column** — every row occupies the same declared number of bytes, and a value outside that range simply cannot be represented — the cost model is a single multiplication: ``` declared_width_in_bytes * row_count = bytes in the value buffer 8 * 10000000 = 80000000 bytes (about 76 MiB) ``` Eighty million bytes. This is not an estimate with error bars around it: for this representation it is the actual size of the packed block of values, because the representation guarantees every row costs the same. ## Why the width does not depend on the values The most common wrong instinct is that a column holding large numbers costs more than one holding small numbers. Under this representation it does not: - The width is fixed **when the column's representation is chosen**, not per value. A row holding 3 and a row holding three billion each occupy the declared width. - There is **no per-value header** — no length, no kind tag, no reference to follow. The buffer is values back to back. - Changing the declared width is therefore the only lever on this figure. Halving it halves the column, and it also halves the range the column can represent. That is the trade. - The row count is the other lever, and it is usually not yours to set. This is also why the figure needs no measurement. You need two numbers — declared width and row count — and both are known before a single value has been read. ## What the figure includes, and what it does not The multiplication describes **the packed value buffer**. A column in practice may carry more beside it, and a table certainly does. | Thing | In the width-times-rows figure? | |---|---| | The packed values themselves | Yes — this is exactly what it counts | | A separate absence mask beside the values | No — a second structure, roughly one bit per row | | The distinct values held once behind a column of codes | No — the figure counts the per-row codes only | | The table's row labels | No — they belong to the table, not to any column | | Anything a stored value refers to elsewhere | No — the buffer holds only the reference | That last row is the one that undoes people, and it is why the arithmetic is reliable for numbers and unreliable for anything separately allocated. ## Where the same multiplication stops being right Two cases, and at this level only two matter: 1. **Text.** Depending on the design, a column of text is either one reference per row pointing at values allocated elsewhere, or one contiguous block of bytes with a position recorded per row. Under the first, width times rows counts the references and misses everything they point at. 2. **Any value reached through a reference.** The same applies to the **holder-for-anything representation** — a column that stores one reference per row and lets each row be a different kind of thing. Its buffer is a slot per row; the cost is elsewhere. In both cases the multiplication still returns a real number. It is simply not the number you wanted: it is the size of the column's own buffer, and the values live somewhere else. ## A worked column list Take a ten-million-row table with four columns: | Column | Representation | Per-row bytes | Value buffer | |---|---|---|---| | identifier | fixed-width numeric | 8 | 80,000,000 | | quantity | fixed-width numeric, narrower | 2 | 20,000,000 | | ratio | fixed-width numeric | 8 | 80,000,000 | | description | text held as references | 8 (the reference) | 80,000,000, plus everything referred to | Three of those four figures are the whole story for their column. The fourth is a floor. Quoting 260 MB for this table would be arithmetic done correctly and an answer that is wrong — the fourth column's real cost is not in it, and neither are the row labels or any absence masks. ## Doing the sizing 1. List the columns and the representation each one carries. 2. For every fixed-width numeric column, multiply declared width by row count. 3. For a column of text or of references, say out loud which storage form it uses; the multiplication is the floor, not the answer. 4. Add the per-column figures, then name what the sum still leaves out. A candidate who does steps 1 and 2 in their head and is honest about step 3 is doing the whole job this question asks for. The trap is stopping after step 2 and calling the result "the table".
- Why does the same arithmetic not survive contact with a column of text?Because under one of the two common storage forms the column's buffer holds a reference per row rather than the value. The multiplication then measures the slots, and every value sits in a separate allocation elsewhere with its own bookkeeping. Under the other form — one contiguous block of bytes plus a position per row — the arithmetic is honest again, but it is characters plus an offset, not a single declared width.
- Does the estimate change if half the values in the column are identical?No. Every row pays the declared width whatever value it holds, and the arithmetic never looks at the values. A representation that stores repeated values as short codes instead is a different representation with a different per-row width, and you would redo the multiplication for that width rather than adjust this one.
- Ten million rows at 8 bytes — is that 80 MB or 76 MB?Both, depending on the unit. It is exactly 80,000,000 bytes, which is 80 MB counting in decimal megabytes and about 76.3 MiB counting in binary ones. The gap is about 5% and it is worth naming when you quote a figure, because the person you hand it to may be reading a tool that uses the other unit.
saying these in an interview costs you the question
- Says a column of large numbers costs more than one of small numbers.
- Cannot give a figure without measuring the stored values first.
- Reports the value buffer as the whole table's size.
- Applies width times rows to a column of text without qualifying it.
- Thinks every row carries its own length or kind tag.