skip to content

A 2 GB text file of records is loaded whole and the process now holds far more than 2 GB. What sets that multiple?

level: juniorimportance: must knowfreq 70%

answer

  1. two encodings, not one transfer
  2. the columns, not the row count
  3. objects behind references against fixed slots
  4. absence mechanism moves the width
  5. the ratio can be under one

basics

~20 s

The in-memory representation of the columns sets the multiple, not the row count. Values held as one runtime object per row, header and reference included, cost several times their text; fixed-width typed columns can cost less than the file did.

solid answer

~50 s

A file and a loaded dataset are two different encodings of the same values, so the file's size is mostly a fact about the encoding it was written in. On load, every column takes whatever representation the tool chose for it. A numeric column parsed out of characters into a fixed-width slot per row often gets *smaller* than its text form. A text column is usually where the multiple comes from: if each value becomes its own runtime object, the column stores a reference per row and you pay an object header and the characters somewhere else. So reason column by column, and expect the ratio to move with the source encoding too - a row-oriented text file loaded into per-value objects can run several times the file, while an already-columnar compact source read into narrow typed buffers, or repeated text kept as small per-row codes, can land near the file size or under it.

go deeper

for a junior

Recall that the file and the loaded data are two different encodings, so the file's size is a weak predictor. Say that text columns are usually where the growth is and numbers often shrink.

for a middle

Explain the mechanics: a per-value object carries a header and the column holds only a reference, so the bytes are scattered and not counted with the column. Contrast that with a fixed-width slot per row.

for a senior

Show that you size from the column mix and a measured sample, and that you know the ratio moves when a representation moves - absence appearing in an integer column, or repeated text losing its per-row codes.

for a principal

The angle is what the team commits to before anyone loads anything: a source encoding and a set of column representations that make the footprint predictable, rather than a rule of thumb that quietly stops holding the first time a column's representation changes.

## Two encodings of the same values A file holds one encoding of the data; a loaded dataset holds another. The file's size is a fact about the encoding it was written in. The **resident footprint** - the bytes the process actually holds - is a fact about the encoding the tool chose on the way in. Nothing forces those two numbers into a fixed relationship, so a remembered multiplier is a guess dressed up as a rule. What is being loaded is usually a **labelled table**: a rectangle of columns where each column carries one type and the rows carry identity of their own. Each column is stored in a **representation** - the fixed in-memory encoding every value of that column is stored in, and the width it costs per row. It is the mix of representations across the columns, not the number of rows, that sets the multiple. The row count scales whatever each row costs; it does not decide what a row costs. ## What the file's size is actually telling you A row-oriented text file spends bytes on things that do not survive loading: - field separators, line endings and quoting characters, which are gone the moment the values are parsed; - numbers written as digits, so a value's cost on disk tracks how many characters it was printed with rather than its magnitude; - repeated text written out in full, once for every row that carries it. And it says nothing about things that only appear once resident: per-value bookkeeping, the structure holding the row labels, and whatever padding or alignment a representation needs. ## Where the growth comes from Growth concentrates in columns whose values are not stored fixed-width inside the column itself: - **Text held as one object per value.** The column stores a reference per row; the characters live elsewhere, each behind its own object header and length. The real cost is what the references point at, and it is scattered rather than contiguous. - **A column holding values of more than one kind**, which usually forces the general per-object representation on the whole column instead of a packed one. - **Absence, in whichever direction the design chose.** Where a design borrows a floating-point sentinel to mean absent, an integer column has no spare bit pattern for it and widens as soon as a hole appears. Where the design records absence in a separate **validity bit** - one bit per row recording present or absent, kept alongside a column of unchanged width - the width does not move at all. Same data, two footprints, and the difference is the absence mechanism rather than the data. ## Where it shrinks Shrinking is just as real, and it is why a flat rule of thumb about the file misleads in both directions: - numbers parsed out of characters into a fixed-width slot per row usually get cheaper; - repeated text stored once in a lookup and referenced per row by a small integer - a **code-per-row representation** - can collapse a wide text column sharply; - an already-columnar, compact source read into narrow typed buffers can land close to its file size, because the file was already storing the values roughly the way memory will. | Source shape | Loaded representation | Direction of the ratio | |---|---|---| | Row-oriented text | one runtime object per text value | several times the file | | Row-oriented text | fixed-width numeric slots | often below the file | | Compact columnar | narrow typed buffers | near one | | Compact columnar | repeated text kept as per-row codes | can be under one | ## How to get an honest number 1. List the columns and what each actually holds - text, numbers, timestamps, mixed values - instead of starting from the file's size at all. 2. For every text-shaped column, establish whether its values are separate objects behind references or one packed buffer with offsets. That single fact usually decides the multiple. 3. Load a sample of rows, read the process's resident size before and after, and take the per-row cost from that measurement rather than from a remembered multiplier. The habit worth building is to stop treating the multiple as a property of the file. Two files of identical size can differ by an order of magnitude once resident, and one file loaded by two tools that chose different representations for its text column will differ too. Reason per column, measure once, and keep the number attached to the representations it came from - when one of those changes, the number is stale.

  • Two files of exactly the same size load into footprints an order of magnitude apart. What would you look at first?
    The column mix, not the files. One is probably mostly numbers, which become fixed-width slots and often shrink, while the other carries free text that becomes one object per value with a header and a reference per row. Check next whether either source was already columnar and compact, because that changes what the loaded form has to build.
  • Does adding rows change the ratio between file size and resident footprint?
    Not by itself. Row count scales both numbers roughly together, so the ratio is stable while the column mix is stable. The ratio moves when a representation moves - a text column that used to be kept as per-row codes falling back to per-value objects, or an integer column widening because absence appeared in it.

saying these in an interview costs you the question

  • Assumes loading is a byte-for-byte transfer of the file into memory
  • Quotes a fixed multiple such as five times the file for any dataset
  • Thinks the row count is what drives the multiple
  • Believes text costs exactly its characters whatever the representation
  • Says a loaded dataset can never be smaller than its file