A packet capture on the interface counts 50k UDP datagrams a minute but your Go agent logs 30k — where did the rest go?
answer
- the capture taps before the socket
- three places a message can vanish
- the kernel counts what it discarded
- the read loop should do nothing else
- a short buffer truncates rather than drops
basics
~20 sAlmost always the kernel's socket receive queue overflowed while the read loop was busy doing other work, so datagrams reached the host and were then discarded. Check the UDP drop counter, then keep the read loop empty.
solid answer
~50 sSeparate three layers before changing anything. a packet capture taps the packet path *before* socket delivery, so a matching capture only proves the datagrams reached the host, not the socket. The kernel counts what it discarded when the receive queue was full — on Linux the drops column in `/proc/net/udp` and the `receive buffer errors` line from `netstat -su`; a climbing counter is the answer. Then instrument the agent itself: count `ReadFrom` calls, bytes and parse errors, because truncation looks completely different from loss — a truncated datagram still counts as a read, so reads match the wire while parse errors climb. The fix is to make the read loop do nothing but read, copy `buf[:n]`, and hand it to a buffered channel that other goroutines drain; parsing, locking and I/O inline in the loop are what let the queue fill. `(*net.UDPConn).SetReadBuffer` then buys headroom for bursts, capped by the system's maximum receive buffer.
code
go · 25 linesvar dropped atomic.Int64
pc, err := net.ListenUDP("udp", &net.UDPAddr{Port: 8125})
if err != nil {
return err
}
if err := pc.SetReadBuffer(4 << 20); err != nil { // ask for 4 MiB of headroom
log.Printf("SetReadBuffer: %v", err)
}
msgs := make(chan []byte, 1024)
buf := make([]byte, 1500)
for {
n, _, err := pc.ReadFromUDP(buf)
if err != nil {
return err
}
msg := make([]byte, n) // buf is reused by the next read
copy(msg, buf[:n])
select {
case msgs <- msg:
default:
dropped.Add(1)
}
}go deeper
Know that a UDP datagram can be lost after it reaches the machine, because the socket has a finite receive queue that fills when the program reads too slowly.
Explain where a packet capture taps relative to socket delivery, and why work done inside the read loop is time the kernel spends filling the queue behind you.
Work the layers in order: wire count, kernel drop counter, in-process counters, then structure. Show that you would move the drop into your own code where it can be counted before you touch any sizing knob.
Own the loss budget. Decide what an acceptable drop rate is for lossy telemetry, what gets alerted on, and when the answer is more capacity or a different transport rather than a bigger buffer.
## Three places a datagram can vanish A datagram that a capture saw but the application never processed died in one of three places, and each has its own evidence: 1. **Before the socket** — wrong port, firewall, checksum error, or the packet was captured on an interface but never delivered. a packet capture sits on the capture path, which is *earlier* than socket delivery, so a clean capture never proves the socket got anything. 2. **In the socket receive queue** — the datagram was delivered, found the queue full, and the kernel discarded it. This is the overwhelmingly common answer for a busy agent. 3. **In your program** — read but then dropped by your own code, or truncated because the buffer was too small, or lost when a full internal channel shed it. Do not skip to a fix before you know which. The three have opposite remedies, and raising a buffer to cure a problem that lives in layer three just adds latency. ## Reading the kernel's own count On Linux the receive-queue drops for a UDP socket are counted per socket in the `drops` column of `/proc/net/udp`, and system-wide in `netstat -su` under `receive buffer errors`. That counter is the cleanest signal in the whole investigation: if it is climbing at roughly the rate of your missing datagrams, the loss is layer two and the cause is that your process did not call `ReadFrom` fast enough. The kernel is telling you the reader is behind, not that the network is bad. If the counter is flat while the numbers still disagree, the loss is inside your program or ahead of the socket, and that is where instrumentation comes in. ## Instrument the agent, not just the host Four counters make the difference visible without a debugger: datagrams read, bytes read, messages successfully parsed, and messages dropped by your own backpressure. Compare them with the wire count: - **Reads far below the wire count, kernel drops climbing** — the read loop is too slow. Layer two. - **Reads matching the wire count but parses far below** — nothing was dropped; datagrams were **truncated** by a short buffer and the parser is rejecting the fragments. Completely different bug, and the byte counts prove it. - **Reads matching, parses matching, but your own drop counter climbing** — the loop is keeping up and the consumers are not. That is a capacity problem downstream, not a socket problem. This is the step people skip, and it is the one that keeps you from turning a sizing knob for a week. ## Fixing the real cause: keep the read loop empty If the loss is in the socket queue, the structural fix is to make the reading goroutine do the minimum possible work per datagram: read, copy, hand off, loop. Every microsecond spent parsing, taking a mutex, appending to an aggregate map or writing to a downstream connection is time the kernel spends filling the queue behind you. ```go n, _, err := pc.ReadFromUDP(buf) msg := make([]byte, n) copy(msg, buf[:n]) select { case msgs <- msg: default: dropped.Add(1) // shed here, where you can count it } ``` Two details matter. The **copy is mandatory** — the loop reuses one buffer, so handing `buf[:n]` to another goroutine hands it an array the next read will overwrite. And the **non-blocking send with a counted drop** converts an invisible kernel discard into a number you own: when you must lose data, lose it somewhere you can see and alert on. ## Then, and only then, size the socket buffer `(*net.UDPConn).SetReadBuffer(bytes)` asks the kernel for a larger receive queue: ```go if err := pc.SetReadBuffer(4 << 20); err != nil { /* log it */ } ``` Understand exactly what that buys. A bigger queue absorbs **bursts** — a spike that arrives faster than you drain, followed by a lull in which you catch up. It buys nothing at all if the average arrival rate exceeds the average drain rate, because a queue that never empties only adds delay before the same drops happen. The request is also capped by the system's maximum receive-buffer setting, and the kernel may adjust the value it grants, so check the returned error and verify rather than assuming you got what you asked for. ## The order to work in Compare the wire count against your read count; read the kernel drop counter; instrument the agent so truncation and downstream loss cannot masquerade as socket loss; empty the read loop; then raise the buffer for burst headroom. Doing it in the other order — buffer first — hides the signal you needed and leaves you with a slower agent that still loses data under sustained load.
- Why doesn't a clean packet capture prove the agent should have received them?The capture path sits before socket delivery, so it shows what arrived at the host, not what reached the socket's receive queue. A datagram can be captured and then discarded because the queue was full, or because nothing was bound to the port. Matching capture and wire counts rules out the network — that is all it does.
- How do the counters distinguish truncation from real loss?Truncation preserves the read count: every datagram still produces one `ReadFrom`, so reads match the wire count while byte counts fall short and parse errors climb. Real loss lowers the read count itself, and the kernel's UDP receive-buffer drop counter rises in step. Reading both numbers separates a buffer that is too small from a reader that is too slow.
- What is the risk of raising the socket receive buffer as the first fix?A bigger queue only absorbs bursts. If the average arrival rate exceeds the rate you drain, the queue stays full and you get the same drops with extra latency in front of them — and you have hidden the drop counter that was pointing at the real problem. Empty the read loop first, then add buffer for burst headroom.
saying these in an interview costs you the question
- Assumes UDP loss must be on the network
- Treats a clean packet capture as proof the app received them
- Raises the socket buffer without reading the drop counter
- Parses and aggregates metrics inline in the read loop
- Confuses truncated datagrams with dropped ones