skip to content

What proves a DHCP binding table is complete before validation drops, and who signs for the ports where a host can still claim the gateway?

level: principalimportance: nice to knowfreq 32%

answer

  1. the table cannot audit itself
  2. compare against sources it did not build
  3. unmatched addresses are the worklist
  4. some hosts never lease at all
  5. an exemption is a signature, not a config line

basics

~20 s

Nothing inside the table proves it. Reconcile against sources it did not build: server leases, switch MAC tables, router neighbour caches, the inventory of hosts that never lease. What stays unmatched becomes a static binding or a signed exemption.

solid answer

~50 s

A binding table records what one switch happened to observe, so it cannot audit itself: four thousand entries prove nothing. Coverage comes from reconciliation against sources the switch did not build: the DHCP server's active leases, switch MAC address tables, the router's ARP and neighbour caches per VLAN, and the asset inventory for appliances that never lease, such as storage controllers and out-of-band management. The number a change board can act on is the count of addresses in those sources with no matching binding, trending toward zero across a window longer than one lease. What remains splits two ways. Static bindings, each owned by whoever owns the appliance and re-certified when it changes. And exemptions, typically an aggregating hypervisor uplink whose guests nobody can enumerate, which means a guest behind it can still claim the gateway for the rack. That is a service owner's signature, not a quiet configuration line.

go deeper

for a junior

Understand that a switch only knows the hosts it observed leasing, so some devices on a segment are simply absent from its table.

for a middle

Be able to name the independent sources that reveal a missing host and explain why each one is partial on its own.

for a senior

Turn reconciliation into a rollout: a per-VLAN unmatched-address count with a trend, a per-rack order, static bindings entered before the flip, and a rollback you have actually tested.

for a principal

Own the exemption. Decide who signs that an unenumerable port stays unvalidated, what compensating position bounds it, and when that decision is reviewed rather than inherited.

## Why the table cannot be its own evidence The question a change board is really asking is "how many hosts will this break?", and the binding table cannot answer it, because it only contains what was learned. A host missing from the table is exactly the host you cannot see in it. So completeness is established by comparison with independent sources, and the whole rollout is designed around producing that comparison. ## The reconciliation sources, and what each one is good for | Source | What it shows | What it misses | | --- | --- | --- | | DHCP server active leases | every address the server believes it issued | anything addressed by hand | | Switch MAC address tables | every device that has sent a frame on a port | no address-to-port authority, only presence | | Router ARP and neighbour caches per VLAN | every address that has answered for itself | includes an address a host claimed but was never given | | Asset inventory / CMDB | the appliances that never lease | is wrong wherever nobody updated it | | Log-only validation failures | the addresses that would have been dropped | proves a mismatch, never intent | Each is partial and two are actively misleading if read alone. Together they bound the population. The useful artefact is a per-VLAN list of addresses seen anywhere with no matching binding, and its trend: it should fall steeply as leases renew, then flatten. The flat residue is the real work. ## The metric to put in front of a board Not "the table is populated". Rather: for VLAN X, N addresses observed with no binding, down from M, unchanged for the last five days, of which so many are identified static appliances with bindings now entered, so many are behind an exempted port, and so many are unidentified. The unidentified count is the honest measure of residual blast radius, and if it is not near zero the change is not ready. Attaching the log-only failure rate over the same window shows the same picture from the control's own side. ## The population you will never learn, and the decision it forces Two groups resist learning. Statically addressed appliances never lease; they need bindings entered from the inventory, and somebody must own that list because it rots the moment an appliance is replaced. And the aggregating uplink — a top-of-rack port carrying a hypervisor link with many guests behind it — cannot be enumerated by the network team at all. That second one is where the organisational decision lives. There are exactly two honest options, and pretending otherwise is how these projects fail: 1. **The guests lease through the switch**, so the table learns them, and the platform team accepts whatever addressing change that requires. 2. **The port stays exempt**, and everything behind it is unvalidated, which means a compromised guest can still answer for the gateway address on behalf of the whole rack. Option 2 is frequently the right call on cost, and it is only defensible when it is written down, owned by the service owner rather than by the network engineer who typed the config, given a compensating position (validation enforced at the boundary the rack sits behind, so the claim cannot spread past it), and given a review date. An exemption with no name attached is the same risk with nobody carrying it. ## What else you are buying at the change call - **A rollout order.** One rack first, chosen because it has a spare path and no storage controllers, not because it was convenient. - **A rollback that is one command per switch**, back to log-only, and a stated trigger for using it. - **A watch window** with the validation drop counters visible, because the failure mode is loss of address resolution and it looks like an application problem from every other angle. - **An outage window** bought from the business, on the understanding that the failure it protects against is gradual rather than instant. ## Where the argument usually goes wrong Engineers arrive with the control and ask for approval; boards approve blast radius and ownership. Bring the unmatched-address count, the named exemptions with their owners, the rollback, and a plain statement of what remains possible after the change — a host behind an exempt port can still claim the gateway for that rack — and the conversation is short. Bring a populated table and a promise, and it is not.

  • The platform team will not enumerate the guests behind the hypervisor uplink. What now?
    Two options, and no third. Either the guests lease through the switch so the table learns them, which costs the platform team an addressing change, or the port stays exempt with a written acceptance from the service owner, a compensating boundary control so a gateway claim cannot spread past the rack, and a review date. What is not acceptable is exempting it silently in a configuration nobody reads.
  • How do you stop the static bindings from rotting after the rollout?
    Tie each binding to the asset record for the appliance it describes, so replacing the appliance is what triggers updating the binding, and keep the validation drop counters visible after the flip. A binding that no longer matches announces itself as drops on that port, which is a far better detector of stale inventory than an annual review.
  • What do you tell the board can go wrong on flip day?
    That hosts without a binding lose address resolution, which looks to their owners like an application failure rather than a network change, and appears within minutes on the affected ports. The mitigation is the per-rack order, the visible drop counters during the watch window, and a one-command rollback to log-only on the switch concerned.

saying these in an interview costs you the question

  • Offers the binding table's entry count as proof of coverage
  • Assumes every host obtains its address by DHCP
  • Treats an exempted uplink as risk-free once it is documented
  • Enforces estate-wide in a single window with no rollback
  • Leaves static bindings with no owner after the rollout

context