Why can an IPv4 traceroute across equal-cost multipath routing report two routers at one hop, or a link between routers that does not exist?
answer
- routers split flows across next hops
- the hash reads header fields
- classic probes change their ports
- hops stitched from different paths
- hold the flow fields constant
basics
~20 sRouters with equal-cost paths pick a next hop per flow by hashing header fields, often including ports. Classic UDP traceroute changes the destination port on every probe, so probes take different paths and the trace mixes routers from several of them.
solid answer
~50 sWith equal-cost multipath, a router has several next hops for one prefix and, as RFC 2991 describes, usually selects one per *flow* by hashing header fields that identify it — which fields is the implementation's choice, and many include the transport ports. Classic UDP traceroute gives every probe a new destination port so it can match replies, which makes every probe a new flow. Probes with the same TTL can then expire on different routers, so one hop lists two addresses, and consecutive hops can come from different branches, drawing a link from hop 3's router to hop 4's that no cable connects. RFC 2991 itself warns that traceroute can give completely wrong results over multiple paths. The fix is to keep the flow-identifying fields identical across all probes of one trace and put the probe identifier somewhere the hash ignores, such as the IPv4 Identification field, which the Time Exceeded quote also returns.
code
pseudocode · 12 lines# Constant-flow trace: every probe hashes onto the same equal-cost path
flow = (src=192.0.2.10, dst=198.51.100.20, proto=UDP, sport=40000, dport=33434)
for ttl in 1 .. max_hops:
for n in 1 .. 3:
probe_id = ttl * 3 + n
send(flow, ttl=ttl, ipv4_identification=probe_id) # flow fields never change
for reply in replies_within(timeout):
quoted = reply.quoted_ip_header
if quoted.identification matches a sent probe_id:
record(hop=ttl_of(quoted.identification), router=reply.source)
if reply is port_unreachable from flow.dst:
stopgo deeper
Recall that routers can split traffic over several equal paths, so a traceroute may show more than one router at the same hop without anything being broken.
Explain per-flow hashing: routers hash header fields to pick a next hop, and classic UDP traceroute changes the destination port every probe, so its probes spread across paths.
Diagnose the artefacts: multiple addresses per hop, impossible links, run-to-run differences. Fix them by holding the flow fields constant, and know per-packet balancing defeats even that.
Treat path visibility as a design property: per-flow hashing gives order and capacity but hides which path a given connection took, so decide what tracing and telemetry the network must offer to make that answerable.
## Equal-cost multipath in one paragraph A router may hold **several next hops of equal cost** for the same destination prefix. Spreading traffic across them adds capacity, but sending consecutive packets of one connection down different paths would reorder them, which hurts TCP. RFC 2991 (Informational) discusses the options: - **per-packet** methods such as round-robin, which reorder traffic and are generally avoided; - **per-flow** methods — modulo-N hashing, hash-threshold (analysed in RFC 2992) and highest random weight — which run a hash over the **header fields that identify a flow** and map the result to one next hop, so every packet of a flow takes the same path. Which fields form the flow is **left to the implementation**. RFC 2991 gives examples ranging from the destination address alone to the triple of source address, destination address and protocol, and warns that including transport ports can be problematic, for instance because non-initial fragments carry none. Many routers nonetheless hash the full five-tuple — addresses, protocol and both ports — as an implementation choice. RFC 2991 also states the consequence for diagnostics directly: "Common debugging utilities such as ping and traceroute are much less reliable in the presence of multiple paths and may even present completely wrong results." ## Why classic traceroute trips over it Traceroute discovers hop *n* by sending a probe with `TTL` *n* and reading the source address of the **ICMP Time Exceeded, type 11, code 0** that comes back. To tell replies apart it needs every probe to be unique within the 8 quoted payload bytes, and classic UDP traceroute does that by **raising the destination port with every probe**. To a hashing router, a new destination port is a new flow. So: 1. Probe A (TTL 3, port p) hashes onto the left branch and expires at router L3. 2. Probe B (TTL 3, port p+1) hashes onto the right branch and expires at router R3. 3. Probe C (TTL 4, port p+2) goes left again and expires at L4; probe D (TTL 4, port p+3) goes right and expires at R4. The printed trace shows **two addresses at hop 3** and two at hop 4. Worse, a reader who joins hop 3's first answer to hop 4's second sees **L3 → R4**, a link that does not exist. Where branches have different lengths, a trace can even appear to loop or to skip a hop. | Symptom in the output | Likely cause | |---|---| | Two or three addresses at the same hop number | probes with different flow fields hashed onto different next hops | | An apparent link between routers that are not adjacent | consecutive hops taken from different branches | | A path that changes between runs | different port choices, or per-packet balancing | ## How to trace through it The principle is to make **all probes of one trace look like one flow**: - keep the source and destination addresses, the protocol and both ports **constant** for every probe; - carry the per-probe identifier in a field that routers do not feed into the flow hash but that still comes back in the Time Exceeded quote — the IPv4 header's **Identification** field is one such place, since the quote includes the whole original IP header; - to map **all** the parallel paths deliberately, run the constant-flow trace several times with different flow values, one per run, and assemble the branches. For ICMP Echo probes the same logic applies with a twist: a router that reads the first four bytes after the IP header as if they were ports sees ICMP's type, code and checksum there, and the checksum changes whenever the sequence number does. A careful tool therefore keeps the checksum constant across probes, compensating in the payload, so every Echo probe looks like the same flow. And if a router balances **per packet** rather than per flow, no choice of fields helps; consecutive probes diverge regardless. ## Related subtleties - The **return path** is chosen independently, by the routers between each hop and you, so even a constant-flow trace measures round-trip times over replies that may take different paths. - **Equal-cost hashing applies to your real traffic too.** Two connections between the same pair of hosts can take different paths, which is why one user's transfer can suffer on a congested branch while another's does not. A trace that should follow that connection must use its addresses, protocol and ports. - Which routes are installed as equal-cost, and how a router selects among route sources, is a routing question of its own; for traceroute the only point is that several next hops exist and a hash picks between them. ## What an interviewer is checking That you can explain the mismatch between **how routers choose a path** (by flow) and **how classic traceroute makes probes unique** (by changing a flow field), predict the symptoms, and name the remedy: hold the flow constant and identify probes by a field the hash ignores.
- A trace that keeps all flow fields constant still shows two routers at hop 4. What might explain it?The router before hop 4 may balance per packet rather than per flow, so identical probes still alternate between next hops; RFC 2991 lists round-robin among the methods. Alternatively the path changed during the trace, which RFC 1393 already noted as a limitation of the hop-by-hop method.
- Why might a TCP trace to port 443 follow a different path than a UDP trace to the same host?Routers that hash the protocol and ports as part of the flow treat the two traces as different flows and may map them to different equal-cost next hops. To follow the path HTTPS traffic takes, the probes must carry that traffic's addresses, protocol and ports.
saying these in an interview costs you the question
- Two addresses at one hop always means the route is flapping
- Equal-cost routing always alternates each packet of a connection between paths
- Which header fields routers hash is fixed by an RFC for all routers
- Changing the destination port per probe has no effect on the path taken
- A constant-flow trace also maps every parallel path in one run