skip to content

Why does declaring a 100 MB zero-initialized static array not add 100 MB to the program image on disk?

level: middleimportance: nice to knowfreq 26%

answer

  1. the file describes, it does not contain
  2. zeros are implied by a size
  3. two static regions, one lifetime
  4. non-zero initial data must ship its bytes
  5. loader supplies a zero-filled range

basics

~20 s

Because zero needs no bytes to record. The image stores the region's size and its symbols, and the loader maps a zero-filled range at start-up. Only static data with non-zero initial contents ships its bytes inside the image.

solid answer

~50 s

Process-lifetime variables are split across two static regions, not one. Data with a **non-zero** initial value must carry those exact bytes inside the program image, because nothing can derive them. Data that starts at **zero** carries only a description - a size and its symbols - because the loader can supply zeros itself. So a 100 MB zeroed array costs a few bytes of description in the file and a 100 MB zero-filled range in the map at load time. This is why file size is a poor proxy for a program's memory footprint in both directions: a small image can map a large zeroed region, and a large image may be mostly instructions. How much physical memory that mapped range eventually costs as the array is touched is an operating-system question with its own owner.

go deeper

for a junior

Remember that a program file describes its memory rather than containing all of it, and a region known to be all zeros needs only its size recorded.

for a middle

Explain the split between the two static regions and what the loader does with each, including why the zero fill is a guarantee and not just an optimisation.

for a senior

Stop anyone reasoning about a service's footprint from its artifact size, and read the mapped regions instead - the two numbers are not related.

for a principal

Weigh shipping a table inside the artifact against building it at start-up: the trade is file size and deploy time against start-up work, not run-time footprint.

## Two static regions, one lifetime Variables that live for the whole process - not per call, not per allocation - are described in the program image and mapped by the loader before the first instruction runs. They land in one of two regions, and which one depends on a single question: **is the starting value all zeros?** | | Initialized static data | Zero-initialized static data | |---|---|---| | Starting value | any non-zero pattern | all zeros | | Carried in the image | the actual bytes | size and symbols only | | Effect on file size | grows by the data size | negligible | | Supplied at load by | copying or mapping the image bytes | a zero-filled range | | Lifetime | the whole process | the whole process | | Writable at run time | yes | yes | Both regions behave identically once the program is running. The distinction exists entirely to keep the file small. ## Why zeros are not worth shipping The image is a description of what the map should look like at entry. For a region whose contents are an arbitrary pattern, the only possible description is the pattern itself. For a region whose contents are all zeros, the pattern is implied by a single fact - *it is zeros* - so recording a size is a complete description. Storing 100 MB of zero bytes in a file and then reading them back would be pure ceremony: the same result, at the cost of 100 MB of storage and the time to transfer it. At start-up the loader therefore does this: 1. Reads the description of each region from the image. 2. Maps the initialized static data so it holds exactly the bytes the image carries. 3. Maps a range of the recorded size for the zero-initialized data, guaranteed to read as zeros. Step 3 is also a correctness guarantee, not only an optimisation: the program may rely on those variables starting at zero, so the range must be zero-filled rather than whatever was previously in that memory. Handing a process another program's leftover bytes would be a serious information leak as well as a bug. ## What it does not make free Three things the trick does *not* buy, worth saying explicitly because they are the follow-up: - **Address space.** The region occupies its full declared size in the map from the moment it is mapped. The array is 100 MB wide as far as addressing is concerned. - **Physical memory as it is used.** Touching the array has a real cost in hardware memory. Exactly how the operating system backs a mapped range and accounts for it is a separate subject with a separate owner; the point here is only that a small file does not imply a small running process. - **A different lifetime.** It is still static data: it exists from load until the process exits, whether or not anything ever reads it. There is no reclamation of a region like this while the program runs. ## Reading a footprint with this in mind This is the reason image size is a bad proxy for memory footprint in both directions. A tiny file may map a very large zeroed region; a large file may be almost entirely instructions and constant tables, whose mapped cost is shared across every process running that program. If you want to know what a process will occupy, read the regions in its map and their sizes - never the size of the file it came from. The same reasoning explains a related observation: a large lookup table filled with meaningful values genuinely does enlarge the image, and is a legitimate size trade to notice. A large table that is computed at start-up and merely declared zero does not, which is sometimes why one design ships a much smaller artifact than another with identical run-time behaviour. ## Where the boundary of the trick is It applies only to storage the program declares at **build** time. A 100 MB block requested at run time was never described in the image at all, so there is nothing there to shrink. The build-time split is precisely what makes the question askable, and it is also why a program that moves a large table from static declaration to run-time construction changes its file size and its start-up behaviour but not necessarily its peak footprint.

  • What changes if that array is given non-zero initial values instead?
    Then the bytes cannot be derived from a description, so they must travel inside the image: the file grows by roughly the size of the data, and start-up has to bring those bytes into the mapped range. The run-time behaviour of the variable is identical - only the file and the load work differ.
  • Does the same trick apply to a large block requested at run time?
    No, because such a block was never part of the image - there is nothing to shrink. The build-time split between initialized and zero-initialized data is a property of how a program describes its static storage, and it simply does not arise for memory the program asks for while running.

saying these in an interview costs you the question

  • Thinks a zeroed static array is not really reserved
  • Says the image must store every byte it will map
  • Believes static data is allocated lazily on first use
  • Uses file size as a proxy for memory footprint
  • Claims zero-filled data is read-only because it ships no bytes