skip to content

Memory Against File Size

A file's size barely predicts what it occupies once loaded, and a run dies at its peak rather than at its resting size. What sets the multiple is what the columns actually hold.

on this pageshow

questions

4

A 2 GB text file of records is loaded whole and the process now holds far more than 2 GB. What sets that multiple?

level: juniorimportance: must knowfreq 70%

answer

  1. two encodings, not one transfer
  2. the columns, not the row count
  3. objects behind references against fixed slots
  4. absence mechanism moves the width
  5. the ratio can be under one

basics

~20 s

The in-memory representation of the columns sets the multiple, not the row count. Values held as one runtime object per row, header and reference included, cost several times their text; fixed-width typed columns can cost less than the file did.

solid answer

~50 s

A file and a loaded dataset are two different encodings of the same values, so the file's size is mostly a fact about the encoding it was written in. On load, every column takes whatever representation the tool chose for it. A numeric column parsed out of characters into a fixed-width slot per row often gets *smaller* than its text form. A text column is usually where the multiple comes from: if each value becomes its own runtime object, the column stores a reference per row and you pay an object header and the characters somewhere else. So reason column by column, and expect the ratio to move with the source encoding too - a row-oriented text file loaded into per-value objects can run several times the file, while an already-columnar compact source read into narrow typed buffers, or repeated text kept as small per-row codes, can land near the file size or under it.

go deeper

for a junior

Recall that the file and the loaded data are two different encodings, so the file's size is a weak predictor. Say that text columns are usually where the growth is and numbers often shrink.

for a middle

Explain the mechanics: a per-value object carries a header and the column holds only a reference, so the bytes are scattered and not counted with the column. Contrast that with a fixed-width slot per row.

for a senior

Show that you size from the column mix and a measured sample, and that you know the ratio moves when a representation moves - absence appearing in an integer column, or repeated text losing its per-row codes.

for a principal

The angle is what the team commits to before anyone loads anything: a source encoding and a set of column representations that make the footprint predictable, rather than a rule of thumb that quietly stops holding the first time a column's representation changes.

## Two encodings of the same values A file holds one encoding of the data; a loaded dataset holds another. The file's size is a fact about the encoding it was written in. The **resident footprint** - the bytes the process actually holds - is a fact about the encoding the tool chose on the way in. Nothing forces those two numbers into a fixed relationship, so a remembered multiplier is a guess dressed up as a rule. What is being loaded is usually a **labelled table**: a rectangle of columns where each column carries one type and the rows carry identity of their own. Each column is stored in a **representation** - the fixed in-memory encoding every value of that column is stored in, and the width it costs per row. It is the mix of representations across the columns, not the number of rows, that sets the multiple. The row count scales whatever each row costs; it does not decide what a row costs. ## What the file's size is actually telling you A row-oriented text file spends bytes on things that do not survive loading: - field separators, line endings and quoting characters, which are gone the moment the values are parsed; - numbers written as digits, so a value's cost on disk tracks how many characters it was printed with rather than its magnitude; - repeated text written out in full, once for every row that carries it. And it says nothing about things that only appear once resident: per-value bookkeeping, the structure holding the row labels, and whatever padding or alignment a representation needs. ## Where the growth comes from Growth concentrates in columns whose values are not stored fixed-width inside the column itself: - **Text held as one object per value.** The column stores a reference per row; the characters live elsewhere, each behind its own object header and length. The real cost is what the references point at, and it is scattered rather than contiguous. - **A column holding values of more than one kind**, which usually forces the general per-object representation on the whole column instead of a packed one. - **Absence, in whichever direction the design chose.** Where a design borrows a floating-point sentinel to mean absent, an integer column has no spare bit pattern for it and widens as soon as a hole appears. Where the design records absence in a separate **validity bit** - one bit per row recording present or absent, kept alongside a column of unchanged width - the width does not move at all. Same data, two footprints, and the difference is the absence mechanism rather than the data. ## Where it shrinks Shrinking is just as real, and it is why a flat rule of thumb about the file misleads in both directions: - numbers parsed out of characters into a fixed-width slot per row usually get cheaper; - repeated text stored once in a lookup and referenced per row by a small integer - a **code-per-row representation** - can collapse a wide text column sharply; - an already-columnar, compact source read into narrow typed buffers can land close to its file size, because the file was already storing the values roughly the way memory will. | Source shape | Loaded representation | Direction of the ratio | |---|---|---| | Row-oriented text | one runtime object per text value | several times the file | | Row-oriented text | fixed-width numeric slots | often below the file | | Compact columnar | narrow typed buffers | near one | | Compact columnar | repeated text kept as per-row codes | can be under one | ## How to get an honest number 1. List the columns and what each actually holds - text, numbers, timestamps, mixed values - instead of starting from the file's size at all. 2. For every text-shaped column, establish whether its values are separate objects behind references or one packed buffer with offsets. That single fact usually decides the multiple. 3. Load a sample of rows, read the process's resident size before and after, and take the per-row cost from that measurement rather than from a remembered multiplier. The habit worth building is to stop treating the multiple as a property of the file. Two files of identical size can differ by an order of magnitude once resident, and one file loaded by two tools that chose different representations for its text column will differ too. Reason per column, measure once, and keep the number attached to the representations it came from - when one of those changes, the number is stale.

  • Two files of exactly the same size load into footprints an order of magnitude apart. What would you look at first?
    The column mix, not the files. One is probably mostly numbers, which become fixed-width slots and often shrink, while the other carries free text that becomes one object per value with a header and a reference per row. Check next whether either source was already columnar and compact, because that changes what the loaded form has to build.
  • Does adding rows change the ratio between file size and resident footprint?
    Not by itself. Row count scales both numbers roughly together, so the ratio is stable while the column mix is stable. The ratio moves when a representation moves - a text column that used to be kept as per-row codes falling back to per-value objects, or an integer column widening because absence appeared in it.

saying these in an interview costs you the question

  • Assumes loading is a byte-for-byte transfer of the file into memory
  • Quotes a fixed multiple such as five times the file for any dataset
  • Thinks the row count is what drives the multiple
  • Believes text costs exactly its characters whatever the representation
  • Says a loaded dataset can never be smaller than its file
open as a page

A step whose finished result occupies 6 GB is killed on a 16 GB machine. What actually had to fit?

level: middleimportance: must knowfreq 58%

basics

~20 s

The peak had to fit, not the result. During the step the input is still referenced while the output and any intermediate are being allocated, so the worst instant can hold several copies of the data at once.

open as a page

A large table is released and live data is now tiny, yet the process's resident size has not dropped. Is that a leak?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Usually not. Resident size is a high-water mark: a general-purpose allocator keeps freed pages for reuse rather than handing them back, so the process stays near its peak with almost nothing live. A leak shows as live bytes that keep climbing.

open as a page

A per-column memory total reports 400 MB for a table while the process holds several gigabytes. What did that total leave out?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A shallow total counts each column's own array of fixed-width slots. Where a slot is a reference, the object it points at - header and characters - is never counted, so text-heavy tables are understated by most of their real cost.

open as a page