skip to content

What does fencing mean when a database standby is promoted, and what mechanisms are used to fence the old primary, including the STONITH (shoot the other node in the head) approach?

level: seniorimportance: must knowfreq 42%

answer

  1. quorum decides, fencing enforces
  2. STONITH: IPMI, PDU, hypervisor stop
  3. SCSI reservation, detach volume, drop VIP
  4. lease expiry + watchdog = self-fence
  5. unconfirmed fence blocks promotion

basics

~20 s

Fencing means guaranteeing the old primary can no longer commit writes or be reached by clients before the new primary opens. It is done by killing the node (STONITH via power or hypervisor), revoking its storage or network access, or removing it from the routing layer.

solid answer

~60 s

Quorum decides who may be primary; fencing enforces that nobody else still is. Without it, an old primary that is partitioned or frozen can keep committing transactions from clients that still reach it. Mechanisms, roughly in order of strength: - **Node-level (STONITH)**: power-cycle or force-off the host via IPMI/BMC, a switched PDU, or a hypervisor/cloud API. Strongest, because a powered-off node writes nothing, and the operation returns a verifiable status. - **Resource-level**: revoke access to the shared resource, for example SCSI persistent reservations on a SAN, detaching a cloud volume, or removing a floating/virtual IP. - **Self-fencing**: the node demotes or shuts down when its leadership lease expires or a hardware watchdog is not petted in time. Cheap and fast, but relies on the sick node behaving. - **Routing-level**: drop the node from the proxy or service registry so clients cannot reach it. Necessary regardless, but weak alone since existing connections may persist. The rule is that promotion must be gated on a confirmed fencing result, not launched in parallel with it. Unfenced promotion is just split-brain with extra steps.

go deeper

for a junior

Know the one-line definition: make sure the old primary truly cannot write before the new one starts accepting writes.

for a middle

Name the main mechanisms (power off, revoke storage or VIP, self-demote on lease loss) and the correct ordering relative to promotion.

for a senior

Discuss confirmation, fence-failure policy, shared-fate fencing paths, watchdog pairing, and the recovery cost of a hard power-off.

for a principal

Set the policy: which failures may be handled automatically, that unconfirmed fencing means stay read-only, and how fencing paths are tested and audited.

## Why fencing exists Election answers who is allowed to write. It does not answer whether someone else is still writing. A primary that is unreachable from the cluster manager may still be alive, healthy from its own point of view, and serving application connections on its side of a partition. If a standby is promoted while that is true, both nodes commit and their histories diverge irreversibly. Fencing is the step that makes the old primary provably unable to commit before the new primary is opened for writes. Its defining property is verifiability: the failover automation must receive a positive confirmation, not merely send a request. ## Mechanisms **Node fencing / STONITH** (shoot the other node in the head). The manager instructs an out-of-band device to remove power or force the machine off: IPMI/BMC, an addressable PDU, a blade chassis controller, or a hypervisor/cloud API call such as a forced instance stop. This is the strongest form because a node without power runs no code. Its weaknesses are practical: the fencing device may itself be unreachable through the same partition, and shared credentials or a single management network can share fate with the failure. Serious deployments configure redundant fencing paths and treat a failed fence as a reason to abort failover. **Resource fencing.** Instead of killing the node, revoke what it needs to write. On shared storage this is a SCSI-3 persistent reservation change or LUN masking; in the cloud it is detaching the managed disk; on the network it is releasing a virtual IP or removing a route. Resource fencing leaves the node running, which is convenient for diagnosis, but you must be certain you revoked every write path, including caches that could flush later. **Self-fencing.** The node removes itself. Two common forms: a leadership lease that must be renewed against a quorum store, where failure to renew triggers immediate demotion or shutdown of the database; and a hardware or software watchdog timer that reboots the host if the agent stops petting it. Self-fencing is fast and needs no external device, and lease-based managers rely on it as their primary defence. It is weaker in principle because it depends on the unhealthy node still executing correct code, which is why a hardware watchdog (which fires even if the agent is wedged) is the recommended pairing. **Routing fencing.** Remove the node from the proxy, service registry, or DNS so new connections cannot reach it. This is necessary in every design, because promotion is worthless if clients keep talking to the old node, but it is not sufficient by itself: already-established connections may survive, and clients with cached endpoints may reconnect directly. ## Ordering and failure handling The safe sequence is: decide leadership, fence and confirm, then promote, then repoint routing. Two consequences follow. First, an unconfirmed fence must block promotion. Many outages are made worse by automation that fired a fence request, ignored the error, and promoted anyway. The correct behaviour when fencing cannot be confirmed is to stay down and page a human, because unavailability is recoverable and divergence often is not. Second, fencing costs recovery time. STONITH implies the old node reboots and must be rebuilt or rewound before rejoining, and a forced power-off skips clean shutdown so crash recovery follows. That is a deliberate trade: seconds of extra downtime against permanent data divergence. ## Practical notes - Fencing devices need their own credentials, monitoring, and periodic testing; an untested fencing path is assumed broken. - A fencing path that traverses the same switch as replication shares fate with the failure it must handle. - Fencing loops (each node fencing the other) are a real hazard in two-node designs and another reason to gate fencing on quorum. - Managed cloud databases hide all of this, but their internals do the same thing: revoke the old writer before opening the new one. ## Interview framing Define fencing as enforcing single-writer, name at least node, resource, and self-fencing, insist on confirmation before promotion, and state the failure policy: if the fence cannot be confirmed, do not promote.

  • The fencing agent cannot reach the old primary's power controller. What should the automation do?
    It should abort the promotion and alert, leaving the cluster unavailable for writes. An unconfirmed fence means the old primary might still be committing, so promoting anyway converts a recoverable outage into unrecoverable divergence. Well-designed managers therefore treat fence failure as fatal, and operators only override manually once they have evidence the node is truly down.
  • How does a watchdog device help when the database host is alive but wedged?
    The cluster agent must periodically pet the watchdog; if it stops, for example because it lost its leadership lease or itself hung, the watchdog resets the host after its timeout without needing any cooperation from software. This bounds the time an unhealthy former primary can keep writing, and it works even when the node is unreachable from outside, which is exactly the case an external fencing device may fail to cover.

Before a new shift takes the controls, you do not just announce it on the radio, you take the old operator's keys and confirm the panel is dark. The announcement is the election; taking the keys is the fence.

saying these in an interview costs you the question

  • Treating removal from the load balancer as sufficient fencing
  • Promoting in parallel with the fencing request instead of after a confirmed result
  • Assuming a node that stopped answering health checks is definitely not writing
  • Configuring a fencing path that shares network, power, or credentials with the failure it must survive
  • Never testing fencing, so it is discovered broken during the first real incident

context