A Linux file server with a 10 Gb NIC never transfers faster than about 90 MB/s and users report occasional stalls. Which `ethtool` and `ip` commands tell you whether the link negotiated the wrong speed or the interface is discarding frames, and how do you read their output?
answer
- do the arithmetic on the throughput first
- speed and duplex before anything else
- errors are physical, dropped is host-side
- rate matters, not the total since boot
- the switch sees the other end
basics
~20 sRun ethtool on the interface to read Speed, Duplex, Auto-negotiation and Link detected; 90 MB/s points at a 1000Mb/s negotiation. Then read ip -s link for RX/TX errors and dropped, and ethtool -S for driver counters.
solid answer
~50 sFirst `ethtool enp3s0`: it prints `Speed:`, `Duplex:`, `Auto-negotiation:` and `Link detected:`. About 90 MB/s is roughly 750 Mb/s, which is a saturated 1 Gb link, so I'd expect to see `Speed: 1000Mb/s` on a card whose `Supported link modes` include 10000baseT/Full — a negotiation or cabling/transceiver problem rather than a software one. Then `ip -s link show dev enp3s0` for the interface counters: **errors** mean malformed frames, which is physical — bad cable, transceiver, duplex mismatch; **dropped** with no errors means frames arrived intact and the host discarded them, which is a host-side capacity problem. `ethtool -S` gives the per-driver breakdown, though the counter names vary by driver. `ethtool -g` shows ring sizes and `ethtool -i` the driver and firmware. I'd also confirm against the switch port, since the host only sees its own end.
code
bash · 13 lines# 1. What did the link actually negotiate?
ethtool enp3s0 | grep -E 'Speed|Duplex|Auto-negotiation|Link detected'
# 2. Are frames being lost, and at what rate?
ip -s link show dev enp3s0
sleep 10
ip -s link show dev enp3s0
# 3. Driver-level detail (counter names vary by driver)
ethtool -S enp3s0 | grep -Ei 'err|drop|miss|crc'
# 4. Ring sizes, in case drops correlate with bursts
ethtool -g enp3s0go deeper
Know that ethtool <interface> prints the negotiated Speed, Duplex and whether a link is detected, and that ip -s link shows per-interface packet, error and drop counters.
Explain the difference between errors and dropped — malformed frames versus intact frames the host discarded — and check counter rates over an interval rather than totals since boot.
Do the throughput arithmetic first to form a hypothesis, then confirm it with the link fields, and know the limits: the host sees one end only, so corroborate with switch-port counters before calling it a cable fault.
Decide how link health is monitored rather than discovered during an incident: which interface counters are collected estate-wide, what thresholds raise a ticket, and how NIC, driver and firmware versions are standardised so negotiation faults are rare.
## Separate the two questions A slow interface has two very different possible causes, and each has its own command: **the link itself negotiated wrong**, or **the link is fine and frames are being lost or discarded**. Doing arithmetic on the reported throughput before touching anything narrows it immediately — 90 MB/s is roughly 750 Mb/s, comfortably inside a gigabit link's real-world ceiling and nowhere near 10 Gb. That is the shape of a link that negotiated at 1 Gb, not of a 10 Gb link losing a few percent. ## What the link negotiated ```bash ethtool enp3s0 # Supported link modes: 1000baseT/Full 10000baseT/Full # Auto-negotiation: on # Speed: 1000Mb/s # Duplex: Full # Link detected: yes ``` The fields that matter: - **`Speed:`** — what the link actually came up at. Compare it with `Supported link modes` and with what the hardware is supposed to be. A 10 Gb card reporting 1000Mb/s is the finding. - **`Duplex:`** — should be `Full`. `Half` on modern equipment is nearly always a negotiation failure, and it produces terrible throughput with collisions. - **`Auto-negotiation:`** — `on` is the correct setting on modern Ethernet. Hard-coding speed and duplex on one end while the other auto-negotiates is a classic way to *create* a duplex mismatch. - **`Link detected:`** — `no` means no carrier at all: cable, transceiver, or the switch port is down or disabled. When the speed is wrong, the causes are physical or configuration-level: a cable or patch panel run that cannot carry the higher rate, a transceiver mismatch, a switch port configured for a fixed lower speed, or a forced setting on the host. `ethtool -i enp3s0` tells you the driver, driver version and firmware version, which matters because firmware bugs affecting negotiation are real and the answer is sometimes a firmware update. ## Whether frames are being lost ```bash ip -s link show dev enp3s0 # RX: bytes packets errors dropped ... # ... # TX: bytes packets errors dropped ... ``` The distinction between the two columns is the whole diagnosis: - **errors** — frames that were malformed on arrival or failed transmission. This is a *physical-layer* signal: a damaged cable, a marginal transceiver, electrical noise, or a duplex mismatch. Errors climbing steadily is a hardware ticket. - **dropped** — frames that were intact but the host did not deliver: no buffer space, the queue was full, or the frame was filtered. This is a *host-side* signal: the receive ring or software queue is too small for the burst rate, the CPU handling the interrupts is saturated, or something upstream in the stack is discarding. Absolute values mean little on a host that has been up for a year. What you want is the *rate*: read the counters, wait, read them again, and see how many were added per second relative to total packets. A handful of errors since boot is noise; hundreds a second is your problem. ## Driver-level detail ```bash ethtool -S enp3s0 ``` This dumps the NIC's own statistics. Two things to say honestly about it: the counter names are **driver-specific**, so what you see on an Intel card differs from a Broadcom or Mellanox one, and the useful ones are found by scanning for names containing `err`, `drop`, `miss`, `discard` or `crc`. CRC-type counters point back at the physical layer; missed or no-buffer counters point at the host not keeping up. Related switches that complete the picture: - **`ethtool -g enp3s0`** — the receive and transmit ring sizes, current and maximum. If drops correlate with bursts and the ring is well below its maximum, enlarging it (`ethtool -G`) is a legitimate mitigation. - **`ethtool -k enp3s0`** — offload features. Offloads being off can raise CPU cost substantially at 10 Gb. - **`ethtool -a enp3s0`** — pause-frame (flow control) settings, which matter when a switch is pushing back. ## The limits of looking only at the host The host only ever sees its own end of the link. A duplex or speed problem is a property of both ends, and error counters on the switch port frequently tell a clearer story than the server's do — the switch may be counting CRC errors on frames the NIC never even handed up. Any confident answer here includes "and I'd have the network team read the switch port counters", because that is what turns a suspicion into a confirmed cable or transceiver replacement. Equally, rule out the boring explanations before blaming hardware: the transfer may be limited by disk, by a single-stream protocol that cannot fill a fat pipe, or by something in the path far from this NIC. The value of `ethtool` here is that it settles the *link* question definitively in one command, so you either have your answer or you have eliminated a whole layer.
- The interface shows `Speed: 10000Mb/s` and `Duplex: Full`, and error and drop counters are flat. Where do you look next?The link layer is exonerated, so the bottleneck is elsewhere: the storage behind the transfer, a single TCP stream that cannot fill a 10 Gb pipe on its own, CPU saturation on the handling core, or something in the path between the hosts. The value of the check is that it eliminated a layer cleanly rather than leaving it as a suspicion.
- Why is `Auto-negotiation: off` with a hard-coded speed usually the wrong configuration?Because negotiation is a two-ended agreement. Forcing speed and duplex on one end while the other auto-negotiates classically produces a duplex mismatch: the link comes up, small transfers work, and throughput collapses under load with rising errors. Modern Ethernet expects auto-negotiation on both ends, so hard-coding is a legacy workaround that tends to cause the failure it was meant to avoid.
- Why can `ethtool -S` output look completely different between two servers?Those counters come from the NIC driver, so the names and the set of statistics are driver-specific — an Intel driver, a Broadcom one and a Mellanox one expose different fields. That is why you scan for substrings like `err`, `drop`, `miss` or `crc` rather than memorising names, and why `ip -s link` is the portable first look.
saying these in an interview costs you the question
- Treating errors and dropped as the same signal
- Reading totals since boot instead of the rate
- Hard-coding speed and duplex to fix negotiation
- Assuming ethtool -S counter names are standard
- Never checking the switch side of the link