skip to content

Why does an HA HDFS cluster have no Secondary NameNode, and what performs checkpointing instead?

level: middleimportance: should knowfreq 44%

answer

  1. a helper, not an heir
  2. the log has to be compacted somehow
  3. snapshot plus write-ahead log
  4. who already has the namespace in RAM?
  5. the standby is doing the work anyway

basics

~20 s

The Secondary NameNode only merges the edit log into a new fsimage so restarts stay fast — it is not a standby. In an HA cluster the standby NameNode already tails every edit, so it performs the checkpoint and the Secondary NameNode role disappears.

solid answer

~50 s

A NameNode persists its namespace as an `fsimage` plus an append-only edit log under `dfs.namenode.name.dir`. Left alone the edit log grows without bound and startup — which replays it — gets slower and slower. **Checkpointing** is the periodic merge of fsimage + edits into a fresh fsimage. In a non-HA cluster a separate `SecondaryNameNode` daemon does that: it pulls the image and edits, merges them, and uploads the new image back. It holds no live namespace and cannot take over. In an HA cluster the **standby NameNode** is already applying every edit from the JournalNodes, so it simply writes out the merged image and ships it to the active — and the Secondary NameNode is not deployed at all. The triggers are the same either way: `dfs.namenode.checkpoint.period` (3600 seconds by default) or `dfs.namenode.checkpoint.txns` (1,000,000 by default), whichever comes first.

code

text · 7 lines
text
# ls of dfs.namenode.name.dir/current
fsimage_0000000000000094321
fsimage_0000000000000094321.md5
edits_0000000000000094322-0000000000000094500
edits_inprogress_0000000000000094501
seen_txid
VERSION

go deeper

for a junior

Recall that the NameNode stores metadata as fsimage plus an edit log, and that the Secondary NameNode merges them — it is not a standby and cannot take over.

for a middle

Explain why the merge exists at all (bounded restart time), who performs it in HA versus non-HA, and the two triggers dfs.namenode.checkpoint.period and dfs.namenode.checkpoint.txns.

for a senior

Show the operational side: alerting on last-checkpoint age, forcing a checkpoint before a planned restart, configuring multiple name directories, and backing fsimage up off-cluster.

for a principal

Frame metadata durability as its own risk class — the blocks are worthless without the namespace — and set the policy for backups, restore drills and acceptable NameNode restart time.

## How the NameNode persists metadata The NameNode keeps the whole namespace in memory, but it must survive a restart. It does that with two artefacts in each directory listed in `dfs.namenode.name.dir`: - **`fsimage`** — a full serialized snapshot of the namespace as of some transaction id. - **`edits`** — an append-only log of every namespace mutation since that snapshot, with the currently open segment named `edits_inprogress_<txid>`. This is the classic snapshot-plus-write-ahead-log pattern. Writes are cheap because they only append to the edit log; recovery is `load fsimage, replay edits`. ## Why checkpointing exists The cost of that design is startup time. If the edit log has been accumulating for weeks, a NameNode restart must replay millions of transactions before it can even enter safe mode, and only then start collecting block reports. On a busy cluster that difference is minutes versus an hour. Checkpointing bounds it: merge the current fsimage with the accumulated edits, write a new fsimage at a higher transaction id, and let the old edit segments be purged. The merge is memory-hungry — it materialises the namespace — which is why it is not done inline on the active NameNode while it is serving traffic. ## Who does the merge: non-HA In a non-HA cluster a separate daemon, the **Secondary NameNode**, does it. On each cycle it asks the active NameNode to roll its edit log, downloads the fsimage and the closed edit segments over HTTP, merges them in its own JVM, and uploads the result back to the NameNode. Because it does exactly this and nothing else, two things follow. First, it needs roughly as much heap as the NameNode. Second — the point interviewers are really testing — **it is not a standby**. It holds no live namespace, receives no block reports, and serves no client RPC. If the NameNode dies you cannot "fail over" to it; the best it offers is a possibly stale image to restore from, and any edits written after its last checkpoint are only in `dfs.namenode.name.dir`. The name is a historical mistake; `CheckpointNode` would have been accurate. (Hadoop also defines `CheckpointNode` and `BackupNode` roles, which are rarely deployed.) ## Who does the merge: HA In an HA cluster the **standby NameNode** is already doing the expensive half of the job. It tails the shared edit log from the JournalNodes and applies every transaction to its own in-memory namespace, so it can serialize a new fsimage at any time and upload it to the active. There is therefore nothing for a Secondary NameNode to do, and running one alongside HA is a configuration error, not a redundancy bonus. If your cluster has multiple standbys, one of them takes the checkpointing duty. ## The knobs Both arrangements use the same triggers, whichever fires first: - `dfs.namenode.checkpoint.period` — seconds between checkpoints, default 3600. - `dfs.namenode.checkpoint.txns` — transactions since the last checkpoint, default 1,000,000. - `dfs.namenode.num.checkpoints.retained` and `dfs.namenode.num.extra.edits.retained` control how much history stays on disk. On a very write-heavy cluster the transaction trigger fires long before the time one; that is intended. ## Protecting the metadata itself Checkpointing keeps restarts fast; it is not a backup strategy. `dfs.namenode.name.dir` accepts a comma-separated list of directories and the NameNode writes the image and edits to all of them, so operators put them on two independent devices (and historically an NFS mount) so a single disk loss does not destroy the namespace. Take real backups of `fsimage` off-cluster as well, because a corrupted or accidentally deleted namespace is the one HDFS failure replication cannot help with — the data blocks on the DataNodes are useless without the metadata that names them. The operational commands worth knowing: `hdfs dfsadmin -safemode enter` followed by `hdfs dfsadmin -saveNamespace` forces an immediate checkpoint on the active NameNode, and `hdfs oiv` reads an fsimage offline (useful for auditing how many files, directories and blocks the namespace holds). ## What to say in an interview Summarise it as: the Secondary NameNode is a log-compaction helper for the non-HA case; HA makes it redundant because the standby already has the namespace in memory; and neither the Secondary NameNode nor checkpointing is a substitute for HA or for backing up the metadata directories.

  • What actually goes wrong if checkpointing stops running for weeks?
    Nothing, until the next restart. The edit log grows unbounded, and the NameNode must replay all of it on startup before entering safe mode, turning a short restart into a very long outage. The metadata directories also fill up because old edit segments cannot be purged. Alert on the age of the last successful checkpoint, not just on daemon liveness.
  • Can you recover a cluster from the Secondary NameNode's copy if the NameNode's disks are lost?
    Only partially, and only to the last checkpoint. The Secondary NameNode's image lags by up to a full checkpoint interval, and every transaction after it lived solely in the lost `dfs.namenode.name.dir`. That means silent, unrecoverable metadata loss. Configure multiple independent name directories and back the fsimage up off-cluster instead of treating this as a plan.
  • How do you force a checkpoint right now, before a planned NameNode restart?
    On the active NameNode, `hdfs dfsadmin -safemode enter` then `hdfs dfsadmin -saveNamespace`, and leave safe mode afterwards. Safe mode is required because the namespace must be quiescent while it is serialized. Doing this before a maintenance restart shortens the edit replay and therefore the downtime.

The fsimage is last month's bank statement and the edit log is every transaction since; checkpointing prints a new statement so you never have to re-add a year of receipts to know your balance.

saying these in an interview costs you the question

  • Calls the Secondary NameNode a backup or hot standby
  • Thinks it serves client reads if the NameNode is down
  • Deploys a Secondary NameNode alongside an HA pair
  • Says checkpointing replicates block data as well as metadata
  • Believes the edit log is truncated automatically without a checkpoint

context