An etcd member backing a Kubernetes cluster starts rejecting writes with "mvcc: database space exceeded". What causes this, and what is the difference between compaction and defragmentation in fixing it?
answer
- MVCC keeps old revisions; quota default 2 GiB
- NOSPACE alarm -> reads ok, writes rejected
- compact = drop history, file size unchanged
- defrag = shrink file, blocking, per member, leader last
- alarm disarm or writes stay blocked
basics
~20 setcd keeps every historical revision and enforces a space quota (2 GiB by default). Compaction discards revisions older than a chosen point, freeing space inside the file; defragmentation rewrites the file so freed pages return to the filesystem. Afterwards you must disarm the NOSPACE alarm before writes resume.
solid answer
~50 setcd is MVCC: each write creates a new revision and old versions are retained until compacted. When the backend database exceeds `--quota-backend-bytes` (default 2 GiB), etcd raises a **NOSPACE alarm** and goes effectively read-only — reads work, every write fails. The two operations differ: - **Compaction** discards history below a revision (`etcdctl compact <rev>`). It frees space *inside* the file; the file does not shrink. kube-apiserver auto-compacts every 5 minutes by default (`--etcd-compaction-interval`), and etcd can self-compact with `--auto-compaction-retention`. - **Defragmentation** (`etcdctl defrag`) rewrites the file to release freed pages back to the filesystem, actually shrinking it. It blocks that member while running, so do it **one member at a time, leader last**. Recovery: compact, defrag each member, then `etcdctl alarm disarm`. Then root-cause the growth — usually Event churn, a hot-looping controller, or a large custom-resource population.
code
bash · 12 linesetcdctl alarm list
etcdctl endpoint status --cluster -w table # compare dbSize vs dbSizeInUse
REV=$(etcdctl endpoint status --write-out=json | jq -r '.[0].Status.header.revision')
etcdctl compact "$REV"
# one member at a time, leader last
etcdctl defrag --endpoints=https://10.0.0.2:2379
etcdctl defrag --endpoints=https://10.0.0.3:2379
etcdctl defrag --endpoints=https://10.0.0.1:2379
etcdctl alarm disarmgo deeper
Recognize that etcd keeps historical revisions, has a size quota, and that exceeding it blocks writes.
State the compaction-versus-defragmentation distinction precisely and know that kube-apiserver compacts every 5 minutes by default.
Walk the runbook end to end — alarm list, compact, defrag one member at a time with leader last, disarm — and then root-cause the churn.
Discuss keeping etcd small as a design constraint: splitting Events onto their own instance, event TTLs, custom-resource volume and object-size budgets, retention versus relist cost, and alert thresholds.
## Why the database grows etcd never updates a key in place. Every mutation appends a new version tagged with a globally increasing revision, and old versions stay readable so clients can watch or read from a past revision. That gives Kubernetes its `resourceVersion` semantics and resumable watches — but it means a workload that rewrites the same object repeatedly grows the store even though the key count is constant. Typical growth drivers: - **Events** — high-churn, short-lived objects; large clusters give Events their own etcd instance via `--etcd-servers-overrides`. - **Hot-looping controllers or operators** patching status every second. - **Lease renewals** from node heartbeats and leader elections. - **Large or numerous custom resources**, or huge ConfigMaps and Secrets (etcd is tuned for small values; the per-value limit is 1.5 MiB and even approaching it is a smell). ## The quota and the alarm etcd enforces `--quota-backend-bytes`, default 2 GiB (values above 8 GiB are discouraged — bigger databases mean slower startup, defrag and restore). On exceeding it, etcd raises a cluster-wide **NOSPACE alarm** and refuses all writes until it is disarmed. This is a safety valve, not a crash: `kubectl get` still succeeds while `kubectl apply` fails with `etcdserver: mvcc: database space exceeded`. Symptoms look like a frozen control plane — no scaling, no rollouts, no new Pods, Leases failing to renew. ## Compaction: dropping history Compaction removes all versions older than a given revision, keeping the latest version of each live key. `etcdctl compact 12345` does it explicitly; more usefully it happens automatically: - **kube-apiserver** compacts on a timer, `--etcd-compaction-interval`, default 5 minutes. In a normal cluster this is what actually keeps history bounded. - **etcd itself** can compact with `--auto-compaction-mode=periodic --auto-compaction-retention=1h` (or `revision` mode). Both compacting is harmless — the later revision simply wins. The visible cost: a watcher resuming from a compacted revision gets `required revision has been compacted`. Kubernetes clients recover by doing a full relist and restarting the watch — the "watch expired, relisting" behaviour informers show. Over-aggressive retention therefore appears as expensive relist storms against the API server. Crucially, **compaction does not shrink the file**. Freed pages are marked reusable inside etcd's backing database, so on-disk size stays at its high-water mark and future writes reuse the space. If you are already at quota, compaction alone may not get you back under it. ## Defragmentation: shrinking the file `etcdctl defrag --endpoints=<one member>` rewrites the backend file, packing live pages together and releasing the tail to the filesystem. Points to state in an interview: - It is a **per-member, local** operation — run it against each endpoint, not once for the cluster. - It is **blocking**: that member stops serving while it rewrites, which for a multi-GB database is seconds to a minute. One member at a time so quorum survives, and do the **leader last** (or move leadership first) to avoid an election mid-defrag. - It is only worth doing when there is a real gap between `dbSize` and `dbSizeInUse` (both shown by `etcdctl endpoint status -w table`). Nightly defrag of a healthy small database is unnecessary churn. ## The recovery runbook 1. Confirm: `etcdctl alarm list` shows NOSPACE; `endpoint status` shows dbSize near quota and dbSizeInUse much smaller. 2. Compact to the current revision (reported by `endpoint status`), or wait for the API server's automatic compaction. 3. Defrag each member in turn, leader last. 4. `etcdctl alarm disarm` — writes do not resume until you do, and this is the step people forget. 5. Root-cause the growth: check Event volume, look for controllers hot-looping on status updates, consider `--etcd-servers-overrides` to split Events onto their own etcd, tune `--event-ttl`, and re-check auto-compaction retention. Raising `--quota-backend-bytes` is a stopgap: it buys time, but a growing database also slows snapshots, restores and startup. ## Preventive posture Alert on `etcd_mvcc_db_total_size_in_bytes` approaching the quota and on the gap between total and in-use size. Watch disk fsync latency (`etcd_disk_wal_fsync_duration_seconds`) too — slow disks and a bloated database produce the same user-visible symptom of a sluggish control plane.
- After a very aggressive compaction, controllers start logging "required revision has been compacted". What is happening and does it matter?A watcher tried to resume from a revision etcd no longer retains, so etcd rejects the watch. Kubernetes informers recover by doing a full relist and restarting the watch, so correctness is preserved. It matters at scale: relist storms are expensive list calls against the API server and etcd, so retention that is too short trades database size for control-plane load.
- Your etcd database keeps growing even though key counts are stable. What would you investigate?Stable key count with growing size points at rewrite churn under MVCC. Look for Events (split them to a separate etcd and shorten --event-ttl), controllers or operators patching status in a hot loop, very frequent Lease renewals, and oversized ConfigMaps or Secrets. Then confirm compaction is actually running by checking the API server's --etcd-compaction-interval and etcd's --auto-compaction-retention.
Compaction is deleting old drafts from a filing cabinet; defragmentation is repacking the remaining folders so you can actually remove a drawer. Deleting alone does not make the cabinet smaller.
saying these in an interview costs you the question
- Saying compaction shrinks the database file on disk
- Running `etcdctl defrag` on all members at once and taking down quorum
- Forgetting `etcdctl alarm disarm`, then concluding the fix did not work
- Treating a raised --quota-backend-bytes as the permanent solution
- Claiming etcd overwrites keys in place, so history cannot be the cause of growth