skip to content

Bundle Analysis

Before you cut anything you have to see what is in the bundle. Being able to describe a concrete analysis workflow separates people who have shrunk a real bundle from people who have only read about it.

on this pageshow

questions

5

In a JavaScript bundle treemap report such as the one webpack-bundle-analyzer produces, what does the area of each rectangle represent, and what does the report not tell you about the code inside it?

level: juniorimportance: should knowfreq 55%

answer

  1. area equals bytes, not time
  2. nested by folder path
  3. three size modes in the report
  4. gzip mode is the network-facing one
  5. shows presence, never execution

basics

~20 s

Each rectangle's area is the byte size that module contributes to an emitted chunk, nested by its folder path. The report shows what is in the bundle and how heavy it is — never whether that code ever runs.

solid answer

~50 s

A bundle treemap is a size map of the build output. Every rectangle is one module — a source file or a file inside a `node_modules` package — and its area is proportional to the bytes that module contributes to the chunk it landed in. Rectangles nest by directory, so a fat `node_modules` block immediately tells you that a dependency, not your own code, dominates. Tools like webpack-bundle-analyzer report three sizes per module: stat size (the raw input), parsed size (after minification), and gzip size (the minified bytes compressed), and only the last is close to what crosses the network. What the map does not say is anything about behaviour: it cannot tell you whether a module is executed, how long it takes to run, or whether it is reachable at all on a given route. It answers "what is in here and how big", which is the first question, not the last.

code

bash · 3 lines
bash
# Produce a webpack stats record from a production build, then render it
npx webpack --mode production --profile --json > stats.json
npx webpack-bundle-analyzer stats.json

go deeper

for a junior

Be able to open a treemap and say what a rectangle means: bytes contributed to the output, grouped by folder. Naming the biggest dependency in a report is the whole ask at this level.

for a middle

Explain the difference between stat, parsed and gzip sizes and which one you would quote to a colleague, and note that the report reflects the build config that produced it.

for a senior

Show that you treat the map as a starting inventory: you check which chunk a heavy module landed in and whether that chunk is on the first-screen path before proposing any change.

for a principal

Own the point that size inventory and runtime cost are different measurements answering different questions, and set the expectation that teams quote a compressed, production-build number when they argue about payload.

## What a bundle treemap actually is When a bundler finishes a build it knows exactly which modules it pulled into each output chunk and how many bytes each contributed. A treemap analyzer takes that record and draws it: the whole canvas is one chunk (or the whole output), and every module inside it becomes a rectangle whose **area is proportional to its byte contribution**. Rectangles are nested by path, so `node_modules/lodash/*` forms one block, your `src/routes/*` forms another. The visual point of a treemap is that size differences that are invisible in a text list — a 400 kB dependency next to two hundred 2 kB source files — become impossible to miss. This is why analysis comes before cutting. Teams that guess at what is heavy almost always guess wrong; they micro-optimise their own components while one date library, one icon set, or one duplicated polyfill accounts for most of the payload. ## Where the data comes from The analyzer is not parsing your shipped files by hand — it reads the build's own statistics record. In webpack that is the JSON stats output; Rollup and Vite expose the equivalent through `rollup-plugin-visualizer`; esbuild writes a metafile that its own online analyzer renders. Two consequences follow: - The report is only as truthful as the build that produced it. A development build has no minification and often extra dev-only code, so a treemap from `pnpm dev` bears little resemblance to production. Always analyze the same command your deploy runs. - Because the record is per chunk, the map also tells you *where* a module landed — main entry, vendor chunk, or a lazily loaded chunk. A dependency being large matters much less when its rectangle sits inside a chunk that only one rarely visited route pulls in. ```bash # webpack: emit the stats record, then render it npx webpack --profile --json > stats.json npx webpack-bundle-analyzer stats.json ``` ## The three sizes, and why the default misleads webpack-bundle-analyzer shows a size selector with three modes, and mixing them up is the most common beginner error: - **stat size** — the size of the module as the bundler read it in, before minification. Useful for seeing how much source a package carries, useless as a user-facing number. - **parsed size** — the module's bytes in the emitted, minified output. This is what the browser must parse and compile. - **gzip size** — the minified bytes after gzip compression. This is the closest available proxy for what actually travels over the network, since servers and CDNs compress text responses. Ranking can change between modes. A package full of long identifiers and comments shrinks dramatically under minification; a package full of large embedded data tables (locale data, emoji tables, country lists) barely shrinks at all and often overtakes it once you switch to gzip. If you make a decision from the default mode without checking, you can spend a week removing something that was already cheap on the wire. ## What the map does not answer A treemap is a static inventory. It says nothing about: - **Execution.** A module can be present and never called. Presence is a size problem; execution is a runtime problem, and they need different measurements. - **Cost in time.** Bytes are correlated with parse and compile time, but a rectangle's area is not a duration. Two equally large modules can differ enormously in how much work they do at startup. - **Removability.** A big rectangle is a *candidate*, not a verdict. Whether it can go depends on who imports it and whether an equivalent lighter path exists. - **What the user downloaded.** The map describes build output. A returning visitor with a warm cache, or a visitor on a route that never loads that chunk, downloads something different. ## Reading one in practice A useful first pass takes about two minutes. Switch to gzip sizes. Look at the top three rectangles by area and name them out loud. Ask, for each, which chunk it is in and whether that chunk is needed for the first screen. Then look for the same package name appearing in two places — a strong hint of a duplicate — and for a package you only meant to use one function from that has arrived whole. Those four observations, in that order, cover most of what a first bundle investigation finds. Everything after that — deciding what to remove, where to split, what limit to enforce — is a separate decision. The map's job is to make sure the decision is about the right module.

  • Your local analyzer report and the deployed bundle disagree on size — what do you check first?
    Whether the report came from the same command the deploy runs. Development builds skip minification and include dev-only branches, and a different environment can flip build-time flags. After that, check compression: the deployed number is usually a compressed transfer size while the report may be showing an uncompressed mode.
  • How would you use these reports to find what a release added?
    Generate a report for the previous commit and the new one from the same build command, then compare them module by module — most analyzers can emit a JSON report alongside the visual one, which makes the diff mechanical. The useful output is a named module or package that appeared or grew, not a total delta.

saying these in an interview costs you the question

  • Says the treemap shows which code is slow
  • Reads stat size as what users download
  • Assumes a big rectangle means the code is unused
  • Treats a large module as automatically removable
  • Analyzes a development build and trusts the numbers

context

open as a page

A bundle report for a production build shows the same package — say date-fns — appearing twice under two different node_modules paths at two different versions. Why does the bundler emit both copies, and how do you get down to one?

level: middleimportance: should knowfreq 50%

basics

~20 s

Two dependents asked for incompatible version ranges, so the package manager installed a nested second copy. The bundler resolves each import to its own file on disk, so both are distinct modules and both ship. You fix it by aligning the ranges, not in the bundler.

open as a page

In a code-split single-page app, how do you work out how much JavaScript a cold first visit to one specific route actually downloads, and why is the size of that route's own chunk a misleading answer on its own?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The honest figure is the sum of every chunk that visit must fetch — runtime, shared and vendor chunks, the route's chunk, and anything it lazily pulls in before rendering. A route chunk's own size counts only the last of those.

open as a page

Your production JavaScript ships as a handful of minified chunks with hashed filenames. How do you attribute those bytes back to the original source files and npm packages, and what do you have to be careful about when you do?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Build the production bundle with source maps and run a source-map-based analyzer such as source-map-explorer over each chunk and its map. It walks every generated byte range back to an original file, giving a size breakdown of what genuinely shipped after minification.

open as a page

Bundle analysis shows one dependency accounts for roughly 40 percent of your main bundle. How do you decide between replacing it, deferring it, or accepting the cost?

level: principalimportance: should knowfreq 36%

basics

~20 s

Decide from what the dependency does for users, not its share of the bundle. Establish who needs it and when, how much of it is genuinely used, what a lighter path would cost in migration risk, and whether removing it would move a metric anyone tracks.

open as a page