Why do `free` and `top` inside a Docker container report the host's memory rather than the container's limit?
answer
- Ask which file the tool actually parses
- A container shares the host kernel
- procfs is global; the process table is not
- The limit is a cgroup attribute
- Look under /sys/fs/cgroup instead
basics
~10 sBecause /proc is not namespaced. free and top read /proc/meminfo, which is host-wide, so a --memory limit is invisible to them. Read the container's cgroup files (/sys/fs/cgroup/memory.max and memory.current) or docker stats instead.
solid answer
~40 sA container is a normal Linux process with namespaces and a cgroup, not a virtual machine, and it mounts the host's procfs. Most of /proc is global: /proc/meminfo, /proc/cpuinfo, /proc/stat and /proc/loadavg describe the machine, not the cgroup. `free` and `top` parse those files, so they print the host's totals no matter what `docker run --memory=512m` set. The limit lives in the cgroup, so read it there: on a cgroup v2 host with the default private cgroup namespace, `/sys/fs/cgroup/memory.max` is the enforced limit and `/sys/fs/cgroup/memory.current` is current usage (on cgroup v1 the names are `memory/memory.limit_in_bytes` and `memory.usage_in_bytes`). From outside, `docker stats` shows the same counters as MEM USAGE / LIMIT. What *is* namespaced is the process table, which is why `top` lists only the container's processes while its summary lines describe the whole machine.
code
bash · 6 lines# On a 128 GiB host, with the container limited to 512 MiB:
docker run --rm --memory=512m alpine free -m
# The number the kernel actually enforces (cgroup v2 host):
docker run --rm --memory=512m alpine \
sh -c 'cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.current'go deeper
Remember the headline: a container shares the host kernel, so free and top show the machine, not your limit. Use docker stats when someone asks how much memory a container is using.
Be ready to explain which parts of /proc are namespaced and which are not, and to name the cgroup files that carry the real limit and usage on both cgroup v1 and v2. Interviewers expect you to say why the limit is still enforced despite the misleading output.
Show the diagnostic habit: when a service is killed but its own metrics look fine, distrust anything sourced from /proc inside the container and go straight to the cgroup counters, separating anonymous memory from page cache before calling it a leak.
Own the fleet-level consequence. Decide whether applications are allowed to auto-size from observed machine capacity at all, whether limits are injected into application configuration as an explicit contract, and whether a compatibility shim like lxcfs is worth its operational cost.
### The confusion this question is testing Engineers who come to containers from virtual machines expect the guest to have its own kernel, its own memory accounting and its own /proc. A container has none of those. It is an ordinary process on the host kernel, wrapped in namespaces (which change what it can *see* and *name*) and placed in a cgroup (which changes what it can *consume*). The reporting tools people reach for first were written for whole machines, and they still behave like it. ### procfs is a view of the kernel, not of your cgroup `/proc` is a virtual filesystem generated by the kernel. Some of it is per-process and therefore genuinely namespaced: the numbered directories `/proc/<pid>` are filtered by the PID namespace, and `/proc/self` resolves inside it. But the machine-wide files are not scoped to a cgroup at all: - `/proc/meminfo` - installed and free RAM for the whole host - `/proc/cpuinfo` - every CPU the host has - `/proc/stat` - host-wide CPU time counters - `/proc/uptime`, `/proc/loadavg`, `/proc/diskstats` - host-wide as well `free` is a thin formatter over `/proc/meminfo`. `top` reads `/proc/meminfo` and `/proc/stat` for its summary lines and the PID directories for its process list. That mixture produces the classic confusing screen: the process list is short and correct (PID namespacing works), while the memory and CPU headers describe a machine with, say, 128 GiB of RAM that your 512 MiB container has never seen. Crucially, none of this means the limit is not applied. Exceeding `--memory` still gets the process group killed by the kernel's OOM killer. The limit is enforced by the memory cgroup controller and simply is not published through procfs. ### Where the real numbers live The limit and the usage are cgroup attributes, so read the cgroup: - cgroup v2: `/sys/fs/cgroup/memory.max` (the enforced limit; the literal string `max` when unlimited) and `/sys/fs/cgroup/memory.current` (bytes charged to the group now). `memory.stat` breaks that down - anon, file (page cache), kernel structures, and so on. - cgroup v1: `/sys/fs/cgroup/memory/memory.limit_in_bytes` and `memory.usage_in_bytes`. An unlimited group reports a very large sentinel number rather than a word, which is its own source of confusion. On a cgroup v2 host, Docker runs containers in a private cgroup namespace by default, so `/sys/fs/cgroup` *inside* the container is the container's own cgroup - the files above are directly readable there. The same files exist on the host under the container's cgroup path, so you can read them without entering the container at all; the exact path depends on the cgroup driver in use. ### Reading it from outside `docker stats` asks the daemon, which reads those same cgroup counters, and prints MEM USAGE / LIMIT and MEM %. Two details matter when you compare it with what the application thinks: 1. If no `--memory` was set, the LIMIT column falls back to the host's total memory - the same misleading number, just displayed honestly as "no limit". 2. Usage is charged to the cgroup including page cache. A process that has read a large file looks fat even though most of that memory is reclaimable. `docker stats` subtracts inactive file cache to get closer to a working-set figure; a raw `memory.current` does not. When usage looks high, check the `file` and `anon` lines in `memory.stat` before concluding there is a leak. ### Why it bites in practice Runtimes and libraries that auto-size themselves from what they observe get this wrong in exactly the same way. Anything that sizes a buffer pool, a cache or a heap from `/proc/meminfo` will pick a number based on the host and then be killed when it grows into it. Container-aware runtimes fixed this by reading the cgroup limit directly instead of procfs, which is the same move you make by hand at the shell. When a runtime is not container-aware, the remedy is to pass the number explicitly rather than to let it guess. ### The workaround: lxcfs If you must make the legacy tools honest - typically for system-container images that run a whole init and a monitoring agent - `lxcfs` is a FUSE filesystem that synthesises cgroup-aware versions of `/proc/meminfo`, `/proc/cpuinfo`, `/proc/stat`, `/proc/uptime`, `/proc/diskstats` and `/proc/swaps`, and you bind-mount them over the container's real ones. Then `free` prints the limit. It is a compatibility shim, not a fix: it costs a host daemon and a FUSE hop, and it does not make the application any more correct than passing it the right number would. For application containers, prefer reading the cgroup. ### The one-line rule Inside a container, /proc answers questions about the machine and /sys/fs/cgroup answers questions about the container. Any measurement that must respect a limit comes from the cgroup or from `docker stats`, never from `free`.
- Why does `top` inside a container list only the container's processes but still show the host's memory and CPU totals?Two different mechanisms. The process list comes from the numbered directories in /proc, which the PID namespace filters, so you only see processes in your own namespace. The summary lines come from /proc/meminfo and /proc/stat, which are machine-wide files with no cgroup awareness. One half of the screen is namespaced and the other half is not.
- `memory.current` looks alarmingly close to `memory.max`, but the service is not being killed. What would you check?How much of that charge is page cache. Usage charged to a memory cgroup includes file-backed pages, which are reclaimable under pressure, so a process that has read large files sits near its limit quite happily. Read `memory.stat` and compare the `anon` figure with `file`. Rising `anon` with a flat workload is the leak signature; a high `file` figure usually is not.
- If a container is running no shell, how would you get the same limit and usage figures?Do not enter the container at all. The cgroup files exist on the host under that container's cgroup path, so you can read them there, and `docker stats` reports the same counters through the daemon. `docker inspect` additionally shows the configured limits, such as HostConfig.Memory and HostConfig.NanoCpus, which is what was requested rather than what is currently used.
It is like reading the building's main electricity meter to find out how much power your apartment is allowed to draw: the meter is real and accurate, it just answers a different question than your own fuse box does.
saying these in an interview costs you the question
- Thinks each container has its own kernel and /proc/meminfo
- Claims free proves the --memory limit was never applied
- Sizes a heap or cache from what free reports in the container
- Assumes all of /proc is namespaced because the PID list is
- Believes docker stats and in-container free should agree
- Reads high memory.current as a leak without checking page cache