A team's service image has grown to 1.8 GB and nobody knows why. Describe how you would find which build steps and which files account for the size, and what you would do with the findings.
answer
- docker history = bytes per instruction
- dive = bytes per file + wasted space
- docker system df -v for real disk
- missing .dockerignore ships .git
- multi-stage before micro-fixes
basics
~20 sUse docker history to attribute bytes to Dockerfile instructions, then dive to see per-layer files and bytes wasted by files deleted or overwritten later. Fix the guilty steps: .dockerignore, multi-stage build, smaller base, same-layer cleanup. Re-measure after.
solid answer
~50 sFirst, per-layer attribution: `docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' myimg:tag`. That maps bytes to instructions and immediately shows whether the base image, a dependency install, or a stray `COPY . .` dominates. Then `dive myimg:tag`, which browses each layer's file tree and reports an efficiency score plus the bytes wasted by files written in one layer and deleted or overwritten in another — exactly what same-layer cleanup would have avoided. `docker system df -v` tells me real daemon disk usage, since summing `docker image ls` sizes double-counts shared layers. Typical findings: no `.dockerignore`, so `.git`, `node_modules` and fixtures got copied in; a full compiler toolchain shipped in the runtime image; a base that should be `-slim` or a JRE. Fixes ranked by payback: `.dockerignore`, then a builder stage copying only the artifact, then base image, then package flags. I rebuild and re-run the same commands to confirm, and add a size check in CI.
code
bash · 6 linesdocker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' myapp:1.4
docker image inspect myapp:1.4 --format '{{.Size}}'
docker system df -v
dive myapp:1.4
CI=true dive myapp:1.4 --lowestEfficiency 0.95 --highestWastedBytes 50MBgo deeper
Know that docker history maps sizes to Dockerfile lines and that .dockerignore prevents accidental copies.
Add dive and the wasted-space concept, and connect each finding to the specific Dockerfile mistake that produced it.
Present a ranked remediation plan, verify by re-measuring, and explain why squashing is a last resort.
Make it systemic: size budgets enforced in CI, shared hardened bases, and an explicit tradeoff between image size, pull latency at scale and patch load.
## Step 1: attribute bytes to instructions `docker history <image>` lists every layer with its size and the command that created it; `--no-trunc` shows the full command and `--format` makes it greppable. Metadata-only layers show 0 B. This answers 'which line did this' fastest. Its limits: it reports layer sizes, not which files inside a layer are large, and history may be missing for images built elsewhere. ## Step 2: attribute bytes to files `dive` is an open-source TUI that loads an image and shows, per layer, files added, modified or removed, with sizes. Two outputs matter: - **Efficiency score / wasted space** — bytes written in one layer and later deleted or overwritten. High waste is the signature of cleanup in a separate RUN, or a file copied then replaced. - **Per-layer file browser** — where you spot the 400 MB of test fixtures or the `.git` directory nobody meant to ship. dive also runs non-interactively (`CI=true dive <image>`) and can fail a build below a configured efficiency or above a wasted-bytes threshold, turning a one-off investigation into a regression guard. ## Step 3: cross-check - `docker image inspect` — total size, layer digests, config. - `docker system df -v` — actual daemon disk usage and what is reclaimable; `docker image ls` counts shared layers once per image, so summing it overstates disk. - `docker save <image> | tar -tv` — raw layer inspection when history is missing. - `docker container diff <container>` — what a running container has written, for when the writable layer rather than the image is what grew. ## Turning findings into fixes Rank by payback: 1. **`.dockerignore`** — near-zero effort. Without it, `COPY . .` ships `.git`, `node_modules`, build output, CI config and sometimes secrets. It also stops unrelated file edits from invalidating the COPY. 2. **Multi-stage build** — compile in a builder stage and copy only the binary, jar or dist into a slim runtime. Usually the difference between 1.8 GB and 200 MB, because compilers, headers and caches never ship. 3. **Right base image** — `-slim` variants, a JRE instead of a JDK, distroless for compiled languages. 4. **Same-layer cleanup and package flags** — the last tens of MB. ## What about squashing? Flattening layers (the legacy experimental `docker build --squash`, or a `FROM scratch` plus `COPY --from` trick) makes deleted files genuinely disappear and can shrink a badly built image fast. The costs are real: layer sharing is lost, so every rebuild is a full transfer rather than a few changed layers; pull-time reuse across images disappears; and build history useful for auditing is gone. Treat squashing as a workaround for an image you cannot restructure, not a strategy. ## Close the loop Rebuild, re-run `docker history` and dive to confirm, and record the size in CI so regressions are visible. Size is not vanity: it drives pull latency during autoscaling, node disk pressure and image garbage collection, and it inflates the vulnerability-scan backlog because every shipped package is something someone must patch.
- docker image ls shows five images at 900 MB each. Is that 4.5 GB of disk?No. Each row counts every layer the image references, and images built from a common base share those blobs on disk. Use docker system df -v for actual usage and what is reclaimable. The same sharing is why pulls are cheap when images share a base.
- Why is image size an operational concern rather than tidiness?Large images slow cold starts and autoscaling because each new node must pull them, they consume node disk and can trigger image garbage collection or eviction pressure, and they enlarge the attack surface and the vulnerability triage backlog because every shipped package is something to patch.
saying these in an interview costs you the question
- Guessing at causes instead of running docker history or dive
- Reaching for --squash before adding .dockerignore or multi-stage
- Summing docker image ls sizes and calling it disk usage
- Assuming docker system prune shrinks an already-built image
- Overlooking that the writable container layer, not the image, may be what grew