In IPv4 over Ethernet, how does ARP find a neighbour's MAC address, and why is the request broadcast but the reply unicast?
answer
- who-has, then is-at
- nobody knows the owner yet
- the request names the requester
- its own EtherType, no IP header
basics
~20 sARP broadcasts a request naming the wanted IPv4 address; the owner answers with a unicast reply carrying its MAC, sent to the requester's MAC copied from the request. The requester caches the mapping, then sends the waiting frame.
solid answer
~50 sBefore a host can put an IPv4 packet into an Ethernet frame it needs the next hop's 48-bit MAC address, and ARP (RFC 826) asks for it. The host broadcasts an ARP request to `FF-FF-FF-FF-FF-FF` in a frame with EtherType `0x0806`, carrying its own MAC and IPv4 address as the sender and the wanted IPv4 address as the target. Broadcast is the only option because the requester does not yet know which station owns that address. Every station on the segment receives it, but normally only the owner replies (a router doing proxy ARP is the deliberate exception), and it can unicast the reply because the request already carried the requester's MAC. The requester caches the mapping and sends its frame; the owner has also recorded the requester's mapping from the request, so its answers usually need no request of their own.
go deeper
Recall the two-step exchange: a broadcast question for an IPv4 address, a unicast answer with the MAC, then the cached mapping and the frame. Say why each direction is broadcast or unicast.
Explain which fields carry the answer, that ARP has its own EtherType and no IP header, and that the owner learns the requester's mapping from the request itself.
Connect the exchange to real symptoms: a lost first packet when resolution is slow, why one broadcast domain bounds ARP, and why unauthenticated replies matter on shared segments.
Weigh ARP's broadcast cost and trust model against IPv6 Neighbor Discovery's multicast design when reasoning about large layer-2 domains and what that implies for segment sizing.
## Why IPv4 needs a resolution step An IPv4 address is a logical address chosen by an administrator or handed out by a DHCP server; an Ethernet station is reached by its **48-bit MAC address**, assigned to the network interface. The two have no arithmetic relationship, so a host that has decided to send a packet to a neighbour still cannot build the Ethernet frame until it knows that neighbour's MAC. RFC 826 (1982) defines the **Address Resolution Protocol (ARP)** to discover the mapping on demand instead of from a hand-maintained table, and RFC 1122 requires IPv4 hosts on Ethernet and IEEE 802 networks to use it. ARP is not carried inside IP. Its packets ride directly in an Ethernet frame whose **EtherType is `0x0806`** (RFC 1042 lists it as decimal 2054), just as IPv4 datagrams use `0x0800`. There is no IP header, no TTL and nothing a router could forward: ARP lives and dies inside one broadcast domain. ## The request: broadcast, because the owner is unknown When the host needs the MAC for a next-hop address it has not cached, it builds an **ARP request** (opcode `1`): - **sender hardware address** - its own MAC; - **sender protocol address** - its own IPv4 address; - **target protocol address** - the IPv4 address it wants resolved; - **target hardware address** - unknown, which is the whole point; RFC 826 leaves its value to the implementation. The frame goes to the Ethernet broadcast address `FF-FF-FF-FF-FF-FF`. There is no alternative: the requester does not know which station owns the address, so it has to ask all of them. Every station in the broadcast domain (on a switched network, every port in the same VLAN) receives the frame and hands it to its ARP module. ## The reply: unicast, because the request named the requester Each receiver compares the target protocol address with its own. Normally only the owner matches; a router configured for proxy ARP, answering for addresses behind it, is the deliberate exception. The owner turns the packet into an **ARP reply** (opcode `2`): it swaps the sender and target fields, puts its own MAC and IPv4 address in the sender fields, and sends it **directly to the requester's MAC**, which it read from the request's sender hardware field. Nobody else needs the answer, so broadcasting it would only make every other station process a packet that is not for it. RFC 5227 notes that broadcast replies are permitted, but it does not recommend them for general use. | | Request | Reply | |---|---|---| | Opcode | `1` | `2` | | Ethernet destination | `FF-FF-FF-FF-FF-FF` (broadcast) | the requester's MAC (unicast) | | Sender fields | requester's MAC and IPv4 address | owner's MAC and IPv4 address | | Target fields | wanted IPv4 address; MAC unknown | requester's MAC and IPv4 address | ## What both sides learn 1. The owner, on receiving the request, records the requester's IPv4-to-MAC mapping **before** it looks at the opcode. RFC 826 assumes communication is bidirectional: if A has a reason to talk to B, B will probably soon talk to A. 2. The requester receives the reply and records the owner's mapping from the reply's sender fields. 3. The requester sends the IPv4 packet that triggered the lookup. RFC 826 suggested simply dropping that packet and relying on a higher layer to retransmit; RFC 1122 section 2.3.2.2 says the link layer SHOULD instead hold at least the latest packet per unresolved address and send it once the address is resolved. 4. Later packets in both directions use the cached entries; how long those entries live is a separate cache-management question. ## A worked exchange Host A is `192.0.2.10` with MAC `00-00-5E-00-53-0A`; host B is `192.0.2.20` with MAC `00-00-5E-00-53-14`; both are on `192.0.2.0/24`. 1. A broadcasts: opcode `1`, sender `00-00-5E-00-53-0A` / `192.0.2.10`, target unknown / `192.0.2.20`. 2. Every station receives it; only B owns the target address. B records `192.0.2.10 -> 00-00-5E-00-53-0A`. 3. B unicasts to `00-00-5E-00-53-0A`: opcode `2`, sender `00-00-5E-00-53-14` / `192.0.2.20`, target `00-00-5E-00-53-0A` / `192.0.2.10`. 4. A records `192.0.2.20 -> 00-00-5E-00-53-14` and transmits its IPv4 packet in a frame addressed to B's MAC. Two frames, one broadcast and one unicast, and both hosts now hold each other's mapping. ## What ARP is not - **Not routed.** A link-layer broadcast does not cross a router, and an ARP packet has no IP header to route on. Off-subnet destinations are reached by resolving the gateway instead. - **Not authenticated.** Any station can send a request or reply with any sender fields, and receivers believe them; forged mappings are an attack topic of their own. - **Not periodic.** RFC 826 explicitly rejects periodic broadcasting of mappings; resolution happens only when traffic needs it. - **Not used by IPv6.** IPv6 replaces ARP with Neighbor Discovery (RFC 4861), which runs over ICMPv6 and sends its solicitation to a multicast group rather than to the broadcast address.
- Can an ARP reply ever be sent as a broadcast?Yes, though it is not the norm. RFC 826 implies replies are unicast, and RFC 5227 says delivering them by broadcast is acceptable but NOT RECOMMENDED for general use, because it doubles ARP broadcast traffic. RFC 3927 (IPv4 link-local addressing) does specify broadcast replies, trading that traffic for faster detection of address conflicts.
- What happens to the IPv4 packet that triggered the ARP request while the host waits for the reply?RFC 826 suggested discarding it and letting a higher layer retransmit. RFC 1122 section 2.3.2.2 tightened that: the link layer SHOULD save at least the latest packet for each unresolved address and transmit it once resolution completes, because otherwise the first packet of every exchange is lost - a TCP connection request or a DNS query then waits for a retransmission.
- Does an ARP request ever reach a host on a different IPv4 subnet?No. ARP is confined to one broadcast domain: routers do not forward link-layer broadcasts, and the packet has no IP header to route. A host sending off-subnet resolves its gateway instead. A router answering for a remote address (proxy ARP) is replying on its own behalf; the request itself still never leaves the segment.
Shouting a name across an open-plan office: everyone hears "who is 192.0.2.20?", but only that person walks over to the one colleague who asked, because the shout said who was asking.
saying these in an interview costs you the question
- The ARP reply is broadcast so every host can update its cache.
- ARP packets travel inside IPv4, like ICMP or UDP.
- Every host that hears the request replies with its own MAC address.
- A router forwards the ARP request to reach hosts on other subnets.
- IPv6 hosts use ARP too, just with longer address fields.