A service sends per-request metrics over UDP to a collector that stalls for thirty seconds; what happens to the datagrams, what does the sender notice, and how should the metric format cope?
answer
- no flow control
- full buffer, silent drop
- send success is local
- cumulative beats delta
- number every datagram
basics
~20 sUDP has no flow control: the sender keeps sending, the collector's receive buffer fills, and further datagrams are dropped silently while the sender notices nothing. Each datagram should carry self-contained, cumulative, sequence-numbered values so loss, duplication and reordering cost precision, not correctness.
solid answer
~50 sNothing in UDP pushes back. RFC 8085 §5 says it plainly: UDP provides no flow control, so the sender does not know whether the receiver can keep up. While the collector is stalled, arriving datagrams queue in its receive buffer until it is full, and after that the host discards them; no message goes back. The sender's send still reports success, because success only means the datagram left the local stack. That is also why metrics often use UDP: a slow collector can never block the service. The format has to absorb it — one complete record per datagram, **cumulative** counters or absolute gauges rather than deltas so a lost or duplicated datagram does not skew totals, a **sequence number** so the collector can drop stale copies and measure loss, and sender-side aggregation to cap the rate.
code
pseudocode · 14 lines# one datagram = one record {sender_id, metric, seq, total}
# sender_id is new on every sender restart, so seq restarts safely
on_datagram(rec):
key = (rec.sender_id, rec.metric)
last = state.get(key)
if last is None:
state[key] = rec # first sight: baseline only
return
if rec.seq <= last.seq:
return # duplicate or late: newer value already held
if rec.seq > last.seq + 1:
lost += rec.seq - last.seq - 1 # gap in numbering = datagrams lost
add_to_series(rec.metric, rec.total - last.total)
state[key] = recgo deeper
Recall that UDP has no flow control, so a slow receiver means dropped datagrams, never a slowed-down sender.
Explain where the drops happen — the receiver's full buffer, or the sender's own interface — and why the send still reports success.
Demonstrate a loss-tolerant format: self-contained records, cumulative values, sequence numbers for dedupe and loss measurement, and sender-side aggregation.
Decide which telemetry may be lossy and which must not, and choose a transport per class of signal rather than forcing one answer on all of them.
## Why UDP gives no backpressure **Flow control** is the mechanism by which a receiver slows a sender down to the pace at which it can consume data. TCP builds it into the protocol: the receiver advertises a window, and when the window reaches zero the sender stops. That is **backpressure** — a slow consumer eventually makes the producer wait. UDP has nothing of the kind. RFC 1122 §4.1.1 lists flow control among the problems a UDP application "must deal directly with", and RFC 8085 §5 is blunter: "UDP provides no flow-control, i.e., the sender at any given time does not know whether the receiver is able to handle incoming transmissions." For telemetry this is usually the *point*. A service emitting metrics must never slow down because its collector is slow, and UDP guarantees it will not. The price is that the collector's problems become silent data loss. ## What happens during a thirty-second stall 1. The collector stops reading — a long pause, a blocked thread, a restart. 2. Datagrams keep arriving and queue in the receive buffer of the collector's host. 3. The buffer fills. RFC 8085 §5 describes this case: when receive calls are issued too infrequently, the receive buffer overflows and data that was sent is not returned. Every further datagram is discarded on arrival. 4. The sender keeps sending at its normal rate. Nothing tells it to stop. 5. The collector resumes and first drains the backlog — values that are now up to thirty seconds old — before it sees current ones. 6. For the rest of the stall's window, the collector has a gap it cannot see unless the format lets it. Loss can happen even earlier: RFC 8085 §5 also notes that datagrams can be lost **inside the sending host** when an application sends faster than the line rate of its outbound interface — for example, every worker flushing at the same instant. ## What the sender can and cannot observe - **Cannot:** that the collector is stalled, that its buffer is full, or that any datagram was dropped. No protocol message reports it. - **Cannot:** infer anything from the send succeeding — success means the datagram was handed to the local stack, not that it was delivered. - **Can, sometimes:** an ICMP Port Unreachable if the collector process is gone entirely and its host is up; RFC 1122 says the host SHOULD send one. A stalled but running collector still holds the port, so no such error appears. - **Can:** whatever the application builds — for example, the collector exporting its own loss count, computed from gaps in sequence numbers. ## Designing the format for the datagram model The format must survive all three of UDP's silent behaviours: loss, duplication and reordering. | Encoding | Lost datagram | Duplicated datagram | Reordered datagrams | |---|---|---|---| | Per-interval **delta** ("+3 requests") | total undercounts forever | total overcounts | harmless for sums | | **Cumulative** counter ("12,345 so far") | next datagram restores the total | harmless — the same total twice | an older value can overwrite a newer one | | Cumulative counter + **sequence number** | gap measured, total restored | discarded | stale value discarded | Rules that follow: - **One self-contained record per datagram.** UDP preserves the boundary, so a record never needs its neighbours to make sense. - **Absolute values beat increments.** Gauges and cumulative counters are idempotent: receiving one twice changes nothing, and missing one is repaired by the next. - **Number each sender's datagrams**, and tie the numbering to the sender's lifetime so a restart starts a fresh sequence instead of looking like stale data. - **Aggregate before sending.** Summarising many events into one datagram per interval caps the rate, reduces loss at both ends and keeps the sender a fair user of the path — RFC 8085 §3.1 expects rate control over all traffic to a destination. - **Stay under the path MTU**, so fragmentation does not turn one lost fragment into a lost record. If some signals must never be lost — billing events, audit records — they do not belong on this path at all; choosing a transport for them is a separate decision.
- Why not have the collector acknowledge each metric datagram so the sender can resend lost ones?Then the sender needs timers, a buffer of unacknowledged records and a policy for a collector that stays down — machinery that lets a stalled receiver slow or bloat the sender, which is what UDP was chosen to avoid. Retransmissions would also fall under the sender's congestion control. If every record must arrive, pick a transport with reliability and flow control; if an estimate will do, loss-tolerant encoding is cheaper.
- Can a UDP metric datagram be lost before it even leaves the sending host?Yes. RFC 8085 §5 notes loss can occur within the sending host when an application sends faster than the line rate of the outbound interface. A burst in which many workers flush at the same moment can be dropped locally, which is another reason to aggregate and spread sends over time.
saying these in an interview costs you the question
- The sender's send call starts failing once the collector's buffer is full.
- UDP slows the sender when the receiver falls behind, just more crudely than TCP.
- Sending counter increments as deltas over UDP keeps the totals exact.
- The collector's UDP layer tells the sender to pause when its buffer fills.
- A running but stalled collector triggers ICMP Port Unreachable at the sender.