A per-column memory total reports 400 MB for a table while the process holds several gigabytes. What did that total leave out?
answer
- two accountings, one name
- slots against what the slots reference
- arbitrary-length values live elsewhere
- packed buffers make the shallow total honest
- record the accounting with the figure
basics
~20 sA shallow total counts each column's own array of fixed-width slots. Where a slot is a reference, the object it points at - header and characters - is never counted, so text-heavy tables are understated by most of their real cost.
solid answer
~50 sThere are two different accountings and they answer different questions. A shallow per-column total sums the width of each column's own slots times its rows. That is honest for a column whose values sit inside the slots - fixed-width numbers, timestamps, per-row codes into a lookup. It is badly wrong for a column whose slots are references, because the characters and the per-object bookkeeping behind them are simply not in the sum. The process's resident size is the other accounting: it counts everything, including what the references point at, the structure holding the row labels, any temporaries still around and whatever the allocator has retained. When the two disagree by an order of magnitude on a text-heavy table, the gap is the objects behind the references, and you should trust the process number for sizing.
go deeper
Recall that a column of arbitrary-length values keeps only a reference per row, and the value itself lives elsewhere. A total that counts columns can therefore miss most of the memory.
Explain which accounting is honest for which representation: slots-only is fine when values sit in the slots, and understates badly when they are references to separate objects.
Show the diagnosis. Take both numbers at the same moment, treat a large gap as a signal about representation rather than as noise, and get a per-column figure by differencing loads if the tool cannot follow references.
The angle is that a footprint quoted without its accounting is not a fact the team can act on. Decide what gets recorded in reviews and capacity conversations so two people's numbers mean the same thing.
## Two accountings of the same table When someone reports what a table occupies, ask which of two numbers they read. The first is a **per-column total**: for each column, the width of its own storage times the number of rows, summed. It is cheap, instant, and attributes cost to a named column, which is exactly what you want when deciding what to change. The second is the **resident footprint** - the bytes the process actually holds. It counts everything the process has, whoever allocated it and whatever it is for. These are not competing estimates of one quantity. They measure different things, and the size of the gap between them is itself diagnostic. ## Why a shallow total misses text A column is an array of fixed-width slots. What is *in* the slot depends on the **representation** - the fixed in-memory encoding every value of that column is stored in: - a fixed-width number, a timestamp or a small integer code sits **inside** the slot, so the slot array is the column and counting slots counts the memory; - a value of arbitrary length usually cannot, so the slot holds a **reference** and the value itself lives elsewhere, behind its own object header, with its own length and characters. A slots-only total counts the references and stops. For a table of a few million rows of free text, that is the small part of the cost, and the total can be off by an order of magnitude while looking entirely plausible. This is a direction-of-design point rather than a universal one. Where a tool keeps a text column as one packed buffer of characters with an array of offsets, or as repeated text held once in a lookup and referenced per row by a small integer, the slots-only accounting is close to honest again, because there is nothing scattered behind a reference. The same shallow number can be trustworthy on one tool and useless on another for the identical data. ## What else sits outside a per-column total Even with no text at all, the two numbers rarely meet: - the structure holding the row labels, and any lookup built to make label matching fast; - temporaries still alive from the step that produced the table; - the process's own baseline - the runtime, loaded code, buffers held by the reader; - pages the allocator has retained after an earlier peak, which belong to the process but to no live value at all. ## Which number answers which question | Question | Number to use | Why | |---|---|---| | Which column should I change? | per-column total, deep where available | it attributes cost to a name | | Will this fit on the machine? | the process's resident size | it counts everything, including what a total omits | | Did my change help? | both, before and after | the per-column total shows intent, the process number shows reality | | What killed the run? | resident size sampled during the step | the worst instant is not visible in either steady number | ## A procedure that settles it 1. Take the per-column total and the process's resident size at the same moment, with the table loaded and nothing else running. 2. If they agree within a modest margin, the columns are sitting in their slots and the shallow total is usable for planning. 3. If they disagree sharply, look for columns whose values are of arbitrary length or of mixed kinds, and check whether the accounting you used follows references or only counts slots. 4. Where the tool offers an accounting that follows references, use it and compare again. Where it does not, load the same table with a subset of columns and read the difference in the process's resident size - that difference is the honest per-column number. 5. Record which accounting the reported figure came from, alongside the figure. A footprint without its accounting is not reproducible by the next person. ## The habit The failure this catches is not innumeracy, it is a category error: two numbers with the same name and different meanings, compared as if one were a rounding of the other. Whenever a footprint is quoted in a review, in a ticket or in a capacity conversation, the useful follow-up is always the same - what did that number count, and what did it follow?
- The per-column total and the process's resident size agree closely. What does that tell you?That the columns are almost certainly sitting in their own slots - fixed-width numbers, timestamps, or repeated text held as per-row codes - with nothing substantial scattered behind references. The shallow total is then usable for planning. It does not tell you anything about the peak during a step, which neither steady number can see.
- How would you get an honest per-column figure when the tool only offers a slots-only accounting?Load the table with the suspect column and again without it, in a fresh process each time, and read the difference in resident size. It is slow and coarse, but it counts whatever the column drags in behind its references, which is exactly what the shallow number omits.
saying these in an interview costs you the question
- Treats a per-column total and the process's resident size as the same measurement
- Says a shallow total is simply wrong, rather than wrong for some representations
- Blames the gap on the runtime's baseline when the table is text-heavy
- Quotes a footprint without saying what the number counted
- Assumes a column's slot always holds the value itself