When an IPv4 server's NIC is replaced and its MAC changes, why can some peers' ARP caches keep it unreachable for minutes, and how do you clear it?
answer
- old MAC, same IPv4 address
- Ethernet has no delivery receipt
- who has heard from the server since
- the merge rule fixes listeners
- static entries never heal
basics
~20 sPeers that cached the old MAC keep sending frames to it, and Ethernet reports no loss. Each recovers only when an ARP packet from the server reaches it, its own ageing or probing drops the entry, or someone flushes it.
solid answer
~50 sA peer's ARP entry still maps the server's IPv4 address to the old MAC, which no NIC owns. Ethernet has no delivery acknowledgement, so the peer sees only timeouts. Recovery follows RFC 826's receive algorithm: a peer that already holds an entry overwrites it from **any** ARP packet the server sends, so a broadcast request from the server fixes every listener that applies that merge, and a request aimed at one peer fixes that peer. That is why the peers the server has contacted since the swap work and the rest do not. Otherwise each peer waits for its own flush: a timeout, from about a minute to hours depending on the implementation, or a failed reachability probe. Flush the entry on affected peers and on the router serving remote clients, replace any static entry, or have the server broadcast; an unsolicited announcement exists for exactly this.
code
pseudocode · 13 lineson receive arp_packet p:
if hardware type or protocol type unsupported: drop
merged = false
if table has entry for p.sender_ip:
entry.mac = p.sender_mac # old MAC overwritten here
merged = true
if p.target_ip is one of my addresses:
if not merged:
table.add(p.sender_ip, p.sender_mac)
if p.opcode == REQUEST:
send REPLY to p.sender_mac
# a peer already holding 192.0.2.20 that hears the server's broadcast
# request for its gateway updates the MAC, though it is not the targetgo deeper
Recall that peers cache the server's old MAC and keep sending to it until their entry is refreshed, flushed or expires.
Explain the merge rule: a peer that holds an entry updates it from any ARP packet the server sends, which is why peers the server contacts recover first.
Diagnose from the pattern: one host unreachable from some peers, no ICMP, recovery after the server initiates traffic. Clear the peers' and the router's entries, fix static ones, and make the server broadcast.
Treat hardware swaps as a change with a network step: announce the new MAC, prefer confirmation-based ageing, and limit static entries to cases whose upkeep someone owns.
## The symptom A server's network card is replaced. It keeps its IPv4 address, but the new card has a new 48-bit MAC address. Afterwards: - some hosts on the same LAN reach it normally; - others time out on every connection, for minutes or longer; - the server itself can reach those failing peers, and once it does, they often start working; - clients beyond the router may fail or succeed as a group. Connectivity to exactly one host is broken, from exactly some places. That pattern points at a cache, and on an IPv4 LAN the cache that maps an address to a MAC is ARP's. ## Why frames to the old MAC vanish silently Each failing peer still holds an entry mapping the server's IPv4 address to the **old** MAC. It builds frames for that MAC and hands them to Ethernet. A switch forwards them to the port where it last saw that MAC, or floods them if that knowledge has aged out; either way no NIC accepts them, because no NIC owns that address any more. Ethernet offers no delivery acknowledgement, so the peer's IP layer learns nothing. No ICMP message is generated: an IPv4 router reports Host Unreachable only when its own ARP request gets no answer, and here nobody is asking. TCP retransmits until it gives up. ## What corrects a peer's entry RFC 826 anticipated this case: "If a host moves, any connections initiated by that host will work", but hosts that connect to it "will have no particular reason to know to discard their old address". What does correct a peer: | Event | Which peers it corrects | |---|---| | the server broadcasts any ARP request, for its gateway or anyone | every peer that already holds an entry and applies RFC 826's merge rule | | the server sends an ARP request whose target is peer P | P, which adds or updates the server as the requester | | P's own flush mechanism fires: timeout, unicast poll, or upper-layer advice | P, on its own schedule | | an administrator flushes P's entry | P, at the next packet | | an unsolicited announcement from the server | every peer that processes it | The first row is the merge rule at work: RFC 826 updates a known sender's hardware address **before** it checks whether the packet was addressed to this host. ## Why some peers and not others 1. **Who has heard from the server.** Peers the server has initiated traffic to since the swap are fixed. A peer that only waits for the server to call it has received nothing. 2. **Whether the server sent any broadcast at all.** A rebooted server usually ARPs for its gateway at once, and that broadcast refreshes every listener. A card swapped without a reboot, while the server's own cache still holds its gateway's entry, may send no ARP packet for a long time. 3. **How each peer ages entries.** An implementation that runs IPv6 Neighbor Discovery's state machine over ARP notices within seconds of sending: STALE, DELAY, unanswered unicast probes, deletion, fresh broadcast. A plain timeout waits it out; RFC 1122 talks of timeouts on the order of a minute, and some routers keep entries for hours as an implementation choice. 4. **Whether the peer accepts unrequested updates.** Some implementations limit RFC 826's merge to harden against forged replies, so an overheard broadcast does not help them. 5. **Static entries.** In common implementations a static entry is never aged and never overwritten. It stays wrong until edited. Clients on other networks reach the server through the router, so they share the **router's** entry: they all fail or all recover together. ## Clearing it and preventing it 1. Confirm the diagnosis: on a failing peer, the server's IPv4 address maps to a MAC that differs from the new card's. 2. Flush that entry on each affected peer, and on the router that delivers traffic from other networks; the next packet triggers a fresh request. 3. Replace any static entry that names the old MAC. 4. Or fix every listener at once by making the server broadcast; an unsolicited ARP announcement, sent whenever an interface's hardware changes, is the purpose-built tool. 5. Afterwards, prefer confirmation-based ageing or moderate lifetimes, and avoid static entries for hardware you expect to replace. ## Misdiagnoses to avoid - Blaming the switch: its MAC table learns the new card the moment the server transmits; the wrong mapping lives in the peers. - Blaming routing or a firewall: those usually show a different pattern, such as a whole subnet failing or an explicit rejection, not one host failing from some peers until the server talks to them. - Waiting for an ICMP error: a wrong cached MAC produces none.
- Why does traffic from a failing peer often start working right after the server contacts that peer?When the server has no entry for the peer, as after a reboot, it broadcasts an ARP request whose target is the peer. Under RFC 826 the target updates or adds the sender's mapping from that request before replying, so the peer now holds the new MAC and its own traffic to the server flows again.
- Why do all clients on other networks fail or recover together?Their packets reach the server's LAN through the router, and the router resolves the server's IPv4 address with its own single ARP entry. While that entry holds the old MAC every remote client is cut off; once it is flushed, aged out, or updated by an ARP packet from the server, all of them recover at once.
- Why does an implementation running the Neighbor Discovery state machine over ARP recover faster than one with a plain timeout?It treats an unconfirmed entry as suspect as soon as traffic is sent: after a short delay without upper-layer confirmation it sends unicast probes to the cached MAC, deletes the entry when they go unanswered, and the next packet broadcasts a fresh request. A plain timeout keeps using the wrong MAC until the timer ends.
saying these in an interview costs you the question
- The switch's MAC table is what holds the server's old address.
- Peers get an ICMP error telling them the MAC changed.
- A static ARP entry will update itself after the next broadcast.
- Peers' ARP caches clear automatically when a neighbour's NIC changes.
- Only rebooting every peer can clear a stale ARP entry.