skip to content

A distributed lock service grants a lease to a client believed to hold exclusive access to a shared resource, such as a storage volume. The client experiences a long GC pause, its lease expires, and a second client acquires the lease and starts writing. The first client then resumes and, unaware its lease expired, also writes. How does fencing prevent this from corrupting the shared resource, and why isn't simply checking the lease is still valid on the client side sufficient?

level: seniorimportance: should knowfreq 50%

answer

  1. zombie client after GC pause/VM stall
  2. fencing token = monotonically increasing number from lock service
  3. resource rejects writes with lower token than high-water mark
  4. enforcement moves from client to resource
  5. STONITH is the hardware-level ancestor of software fencing

basics

~20 s

Fencing gives each lease a rising number; the resource itself refuses any write whose number is older than the newest it has seen. That way even a confused old client that thinks it still owns the lock physically can't write, because the resource, not the confused client, enforces the check.

solid answer

~50 s

Fencing solves the zombie-client problem inherent to any lease scheme: a client can't reliably know, from its own side, whether it's still the legitimate holder, because pauses (GC, VM suspension, disk stalls) can make it fall behind reality without realizing it. The fix is to have the lock service hand out a monotonically increasing fencing token with each lease grant, and require the protected resource itself to reject any write carrying a token lower than the highest it has already seen. This moves enforcement from the unreliable, possibly paused client to the resource, which can compare tokens cheaply and correctly regardless of how confused or delayed any given client is. Checking lease validity client-side isn't sufficient because the check-then-act sequence has a gap: a pause can occur between the validity check and the actual write, and the client has no way to know time passed while it was frozen.

go deeper

for a junior

Understands, at a basic level, that a paused client can wrongly believe it still owns a lock, and that some mechanism is needed to stop it from corrupting data.

for a middle

Can describe the fencing token mechanism, an increasing number the resource compares against a high-water mark, and why it's attached to every write.

for a senior

Can explain precisely why client-side re-validation fails, the check-then-act gap under unbounded pauses, and why enforcement must live at the resource boundary, not upstream of it.

for a principal

Can assess whether a given storage layer is fencing-capable, design a conditional-write-based fencing scheme for one that isn't, and connect the technique to its hardware-era ancestor STONITH when justifying the design to a team.

## The gap fencing closes **Fencing** is the mechanism that closes a specific, dangerous gap in lease- and lock-based coordination: the gap between a client believing it holds exclusive access and that belief actually still being true. Distributed locks (via ZooKeeper, etcd, Chubby, or a database-backed lease table) are typically implemented as time-bounded **leases** - a client is granted ownership for, say, 10 seconds, must periodically renew it, and the lock service considers the lease expired and re-grantable if renewal doesn't arrive in time. The problem is that a client can be paused for an unbounded, unpredictable amount of time: - a stop-the-world garbage collection pause - a hypervisor migrating the VM - an overloaded kernel scheduler - a slow disk I/O syscall blocking a thread During such a pause it cannot renew its lease, has no way to observe that time has passed, and has no way to know, once it resumes, whether its lease already expired and was reassigned. From the paused client's own perspective, no time appeared to pass at all; it resumes execution exactly where it left off, fully believing it's still the lock holder, and proceeds to write to the shared resource. Meanwhile the lock service, having seen no renewal, has already reassigned the lease to a second client, which is now also writing. Both clients now believe, correctly by their own local reasoning, that they are the exclusive owner - this is the **zombie** or **split-brain writer** problem. ## Why client-side revalidation falls short Client-side revalidation - having the client re-check whether its lease is still valid right before writing - looks like an obvious fix but doesn't actually close the gap, because the check and the subsequent write are two separate steps with a time window between them, and the pause that caused the original problem can just as easily land in that window. - A **check-then-act** sequence is never atomic against an unbounded external pause; you can always construct a schedule where the pause happens after the check succeeds but before the write lands. - No amount of checking harder or more recently on the client side eliminates this, because the client cannot make guarantees about its own future scheduling. ## The token the resource enforces The fix fencing provides is to move the correctness check out of the untrustworthy, pausable client and into the resource being protected, where it can be made a real invariant rather than a best-effort courtesy check. Concretely: - Every time the lock service grants or renews a lease, it issues a **fencing token** - a number that increases monotonically across every grant, typically a simple counter maintained by the lock service itself. - The client must attach this token to every write it sends to the protected resource. - The resource - the storage node, the database, the file server - is modified to track the highest fencing token it has ever accepted and to reject any incoming write whose token is lower than that **high-water mark**. In the scenario above: 1. Client A gets token 33, pauses, its lease expires. 2. Client B gets token 34 and writes successfully, and the resource's high-water mark becomes 34. 3. Client A resumes and sends its write tagged with token 33, and the resource, which now requires tokens 34 or higher, rejects it outright, regardless of how confident client A is that it still holds the lock. The resource doesn't need to know or care why token 33 is stale; it just enforces **monotonicity**, which is a property it can check locally and correctly without needing to reason about clocks, pauses, or leases at all. ## Safety moves to the last-writer boundary This is a meaningful architectural shift: fencing pushes safety enforcement to the last-writer boundary - the resource - rather than relying on any actor upstream of it (client, lock service, network) behaving correctly or promptly. It's the same principle behind fencing in the older, more literal sense from cluster and SAN environments - **STONITH** (shoot the other node in the head) forcibly powers off or disconnects a suspect node from shared storage at the hardware level, because you can't trust a possibly-still-running zombie node to voluntarily stop. Software fencing tokens achieve the same isolation logically, without needing to physically kill anything, by making the shared resource itself the arbiter. ## The trade-off The trade-off is that fencing only works if the protected resource is willing and able to enforce the check - it requires that resource to be **fencing-aware**: - maintaining per-key or per-resource high-water marks - comparing tokens on every write which not every storage system supports natively. Plain object stores don't have a built-in fencing-token check, so applications must layer one on, for example via conditional writes keyed on a version or `ETag` field. If the resource can't enforce token ordering, fencing degrades back to a best-effort convention that a sufficiently pathological zombie client could still violate. This is why systems that need strong lock-based mutual exclusion design the storage layer itself to be fencing-aware rather than trusting the lock service or client discipline alone.

  • Why can't the lock service itself just forcibly notify the paused client that its lease expired?
    A paused client, by definition, isn't executing and can't receive or process any notification during the pause - that's exactly the failure mode. Any notify-the-old-holder approach still has to deal with the resumed client acting on stale beliefs once it wakes up, which is the same gap fencing tokens are designed to close at the resource, not the client.
  • How does a fencing token differ from simply checking a timestamp or lease expiry time on the write?
    A timestamp comparison depends on clocks being synchronized and doesn't have a hard, resource-enforced high-water mark - it's still a value the client computes and attaches based on its own possibly stale belief about time. A fencing token is assigned centrally by the lock service as a strictly increasing counter and compared against a value the resource itself remembers seeing, making the check independent of wall-clock drift or the client's confidence about elapsed time.
  • What happens if the protected resource is a plain object store that has no native concept of a fencing token?
    You have to layer the equivalent behavior on top, typically via conditional writes such as compare-and-swap semantics keyed on an ETag or version field stored alongside the object, so the last accepted version acts as the high-water mark. Without any such conditional mechanism, fencing can't be enforced at that resource, and a zombie client's write could silently succeed and corrupt data.

It's like a hotel that doesn't trust guests to know their keycard was deactivated when their room was reassigned - instead the door itself only accepts the newest keycard issued for that room, so even a confused guest with an old key physically cannot get in, no matter how sure they are it's still their room.

saying these in an interview costs you the question

  • Proposes fixing the zombie-client problem by having the client re-check its lease right before writing, without noting the check-then-act gap
  • Doesn't realize the resource, not the client or lock service, must enforce the token check
  • Confuses fencing tokens with simple lease-expiry timestamps
  • Assumes a paused client can be relied upon to notice time has passed
  • Thinks fencing eliminates the need for leases/locks entirely rather than complementing them

context