You run `mtr` toward a server and hop 6 reports 40% loss while the final hop reports 0%. What does that intermediate loss tell you, and what pattern in an mtr or traceroute report indicates loss that is actually hurting your traffic?
answer
- the middle rows measure something else
- control-plane work, not forwarding
- routers police the replies they generate
- loss must persist to the last row
- forward path only, return path invisible
basics
~20 sLoss at a middle hop that does not persist to the hops after it is almost always an artifact: that router deprioritises or rate-limits the replies it generates, while still forwarding traffic fine. Real loss shows at a hop and every hop beyond it, including the destination.
solid answer
~50 sIntermediate hops in `mtr` only appear because they generate an ICMP time-exceeded message when a probe's TTL runs out, and generating that message is control-plane work that routers deliberately rate-limit and deprioritise. So a router can drop 40% of the replies it owes *you* while forwarding 100% of the packets that pass *through* it. The rule of thumb is that loss must be cumulative to be real: if hop 6 shows 40% but hops 7, 8 and the destination all show 0%, hop 6 is lying and your traffic is fine. If hop 6 shows 40% and every hop after it shows roughly 40% or more, you have found where the path genuinely degrades. Two more caveats: mtr shows you the forward path only, so an asymmetric return path can be the real culprit, and probes to a service port with `mtr --tcp --port 443` are treated more like real traffic than the default probes are.
code
bash · 1 linemtr -rwc 100 --tcp --port 443 api.example.comgo deeper
Know that mtr combines traceroute and ping, and that the destination row is the one that tells you whether your traffic is getting through.
Explain why intermediate hops under-report: those replies are generated by the router's control plane and are rate-limited, so their loss figure is not a forwarding statistic. State the cumulative rule explicitly.
Demonstrate that you can turn a report into an action — sampling long enough to be sure, probing the production port, checking from a second vantage point, and knowing when the finding belongs to a provider rather than to you.
Own the escalation and design angle: what evidence a transit provider will actually accept, and whether recurring single-path degradation justifies multi-homing or a different egress rather than repeated tickets.
## Where the numbers in an mtr report come from `mtr` is traceroute and ping fused together. It sends probes with deliberately small TTL values: a probe with TTL 1 expires at the first router, TTL 2 at the second, and so on. Each router whose TTL counter hits zero is expected to send back an ICMP time-exceeded message, and that reply is what identifies the hop. Probes addressed all the way to the destination are answered by the destination itself. mtr repeats this continuously and builds a table with `Loss%`, `Snt`, `Last`, `Avg`, `Best`, `Wrst` and `StDev` per hop. The crucial asymmetry hides in that description. For an intermediate hop, the number you are measuring is *how reliably that router talks to you about itself*. For the destination row, you are measuring *how reliably packets get all the way there and back*. Those are different questions, and only the second one is about your traffic. ## Why middle hops lie Forwarding a packet is fast-path work: on a real router it happens in hardware or a tight dataplane loop. Generating an ICMP time-exceeded message is control-plane work handled by the router's CPU, and every serious vendor rate-limits it, precisely so that a flood of expiring packets cannot become a denial of service against the router's brain. Some operators deprioritise or drop these replies entirely. The result is the single most misread output in network troubleshooting: a scary red row in the middle of an mtr report from a device that is passing your traffic perfectly. ## The cumulative rule Read an mtr report top to bottom and ask whether loss *persists*: ``` 5. router-a.isp.net 0.0% 100 ... 6. router-b.isp.net 40.0% 100 ... <- artifact 7. router-c.isp.net 0.0% 100 ... 8. api.example.com 0.0% 100 ... <- what you actually care about ``` Hop 6 cannot be dropping 40% of transit traffic while hop 7 and the destination see none — the packets that reach hop 7 went *through* hop 6. That row is ICMP policing, nothing more. The pattern that matters is the opposite one: loss that appears at some hop and continues, at roughly the same level or worse, through every hop after it and at the destination. That is the point where the path starts dropping packets. In practice the last row is the row that decides whether you have a problem at all; the hops above it only help you locate it. ## The asymmetry caveat Every hop you see is a hop on the *forward* path. The replies take whatever return route the internet chooses, and that route is frequently different. So a hop that shows genuine loss may be innocent, with the damage happening on a return path you cannot see from this end. When it matters, run mtr from both ends and compare. ## Making the probes look like real traffic Default Linux traceroute sends UDP datagrams to high ports; mtr defaults to ICMP echo. Both are shaped and filtered differently from your production traffic. When you suspect policy rather than congestion, probe the real port: ```bash mtr -rwc 100 --tcp --port 443 api.example.com traceroute -T -p 443 api.example.com tracepath api.example.com ``` `mtr -rwc 100` is the form you paste into a ticket: `-r` produces a one-shot report, `-w` uses wide output so full hostnames fit, and `-c 100` fixes the sample count so the loss percentages mean something. `traceroute -T` sends TCP SYNs and needs privileges; `tracepath` needs none. ## What a good answer sounds like Name the mechanism (control-plane ICMP generation is rate-limited), state the cumulative rule, and then say what you would do with the finding. Even genuine mid-path loss inside someone else's network is usually not yours to fix — the useful outcomes are evidence for the provider, or a decision to route around it. An interviewer is checking that you do not open a ticket every time one row in an mtr report turns red.
- Suppose loss does persist from hop 6 through to the destination — what do you do with that?Confirm it is stable rather than a momentary blip by running a longer sample, and run mtr from a second vantage point to see whether the same hop is implicated. Then identify who owns that hop from its hostname and address, and hand them the report. If it is a transit provider you do not control, the practical outcomes are an escalation with evidence or a routing change, not a local fix.
- Why might mtr with default probes show a clean path while your application still sees losses and stalls?Default probes are small ICMP or UDP packets, treated differently from a real TCP flow: they are not subject to the same queueing, they never grow to full-size segments, and they can be rate-shaped separately. Congestion that only appears at production packet sizes and rates, or a policer that targets one port, will not show up. Probe with --tcp on the real port, and compare against the application's own retransmission behaviour.
- What does the StDev column in an mtr report add over Avg?It quantifies jitter. A hop with a modest average but a large standard deviation is queueing unevenly, which hurts latency-sensitive traffic even when no packets are lost at all. Averages alone hide that, so a stable Avg with a rising StDev is a real signal worth chasing.
Asking each router about its own health is like asking people along a relay route to shout their name as your parcel passes. A tired shouter proves nothing about whether the parcel arrived; only the recipient's answer does.
saying these in an interview costs you the question
- A red middle hop means that router is dropping my traffic
- Every hop's loss percentage is independent evidence
- Higher latency at one hop proves that hop is congested
- mtr shows both directions of the path
- ICMP loss and TCP loss are always the same number