With 1-in-1000 packet sampling in NetFlow, IPFIX or sFlow at a 100 Gb/s edge, how do you scale sampled counts up, and how far can you trust them?
answer
- multiply by the sampling rate
- the collector must learn N
- error follows the sample count
- roughly one over root c
- lost samples lower the attained rate
basics
~20 sMultiply sampled packet and byte counts by N, the 1-in-N sampling rate. With random sampling, an estimate from c samples has a relative error near 1/sqrt(c): big aggregates are accurate to a fraction of a percent, small flows are mostly noise.
solid answer
~40 sRFC 5474 renormalises sampled volumes by dividing by the selection fraction, which for 1-in-N sampling means multiplying by N. The collector learns N from the export: NetFlow v9's `SAMPLING_INTERVAL` in options data, IPFIX's sampling Information Elements from RFC 5476, or the `sampling_rate` in every sFlow v5 flow sample. The estimate's precision depends on how many samples it rests on, not on N alone. With random selection, c samples give a relative standard error of about 1/sqrt(c), roughly ±1.96/sqrt(c) at 95%. At 100 Gb/s with 800-byte packets, that is about 937,500 samples a minute for the whole link, within about ±0.2%. A prefix with 40 samples scales to about 40,000 packets, ±31%. Lost samples reduce the attained rate, so correct with sFlow's `sample_pool` or sequence numbers.
go deeper
Recall that sampled counts must be multiplied by the sampling rate and that the result is an estimate, not a total.
Explain where the collector learns N in v9, IPFIX and sFlow, and why lost samples lower the effective rate.
Compute an estimate's error from its sample count, judge which aggregates are trustworthy at 1-in-N, and correct for drops using sample_pool or sequence numbers.
Choose sampling rates per edge from the smallest aggregate that must be accurate, balancing exporter and collector load against the error each report can carry.
## Why an edge samples at all At 100 Gb/s with an average packet of 800 bytes, a link carries 100,000,000,000 / (800 x 8) = **15,625,000 packets per second**. Building flow state for every one of them is expensive, so edges commonly **sample**. Each of the three formats does it differently: - **sFlow v5**, an sflow.org specification and not an RFC (RFC 3176, Informational, describes the older version 4), copies the header of roughly 1 packet in N and sends each sample to the collector. - **Sampled NetFlow v9** (RFC 3954, Informational) and **IPFIX** (RFC 7011) feed only the selected packets into the flow cache, so a record's counters cover sampled packets only. - The IETF's packet-sampling documents are **PSAMP**: RFC 5474 is the Informational framework, and RFC 5475 and RFC 5476 are the Standards Track selection techniques and their IPFIX export. At 1-in-1000 the link above yields 15,625 samples a second, or **937,500 a minute**. ## Scaling up RFC 5474 §11.1 says absolute volumes are estimated "by renormalizing the sampled traffic volumes through division by" the selection fraction. For 1-in-N sampling that means multiplying: - estimated packets = sampled packets x N; - estimated bytes = sum of sampled packet lengths x N. The collector must know N, and each format exports it: | Format | Where the sampling rate travels | |---|---| | NetFlow v9 | options data: `SAMPLING_INTERVAL` (type 34), `SAMPLING_ALGORITHM` (type 35) | | IPFIX / PSAMP | options data with Information Elements such as `samplingPacketInterval` (RFC 5476) | | sFlow v5 | `sampling_rate` and `sample_pool` in every `flow_sample` | The **configured** fraction is not always the one achieved. RFC 5474 defines an **attained selection fraction** computed from reported sequence numbers, which covers both sampling and transport loss. sFlow carries `sample_pool`, the packets that could have been sampled, and a `drops` counter for samples the agent could not process. The sFlow text puts it plainly: lost flow samples mean "a slight reduction in the effective sampling rate". A collector that multiplies by the configured N after losing 20% of datagrams reports 20% low. ## How much to trust the estimate With **random** selection, each packet is chosen independently with probability p = 1/N. If the true count is n, the number sampled is binomial with mean c = n/N. The scaled estimate has a relative standard error of sqrt((1 - p) / c), which is about **1/sqrt(c)** when N is large. At 95% confidence, use roughly **±1.96/sqrt(c)**. | What is measured | Samples c | 95% relative error | |---|---|---| | whole 100 Gb/s link, one minute | 937,500 | about ±0.2% | | a prefix, 1-in-100, one minute | 400 | about ±9.8% | | a prefix, 1-in-1000, one minute | 40 | about ±31% | | a small flow | 3 | about ±113% | The error depends on **the number of samples**, so: 1. Large aggregates, such as link totals, top destination ASes or busy prefixes, are trustworthy. 2. Small flows are not. Their estimates are a few multiples of N, and many small flows get no sample at all. 3. To reach a target precision, collect enough samples. ±10% at 95% needs (1.96 / 0.10)^2, about **384 samples**. Sample more often or aggregate over a longer window. A larger N reduces samples and widens the error. Byte estimates carry extra variance because sampled packets differ in size, so treat the packet-count figure as the better case. ## Random versus deterministic selection RFC 5475 §5.1 warns that **systematic** sampling, every Nth packet, "always involves the risk of biasing the results" when the traffic itself has a period. NetFlow v9 labels the method in `SAMPLING_ALGORITHM` as deterministic (0x01) or random (0x02). sFlow v5 requires that every packet has "an equal chance of being sampled" and resets its skip counter to a random value after each sample. ## A worked estimate ```pseudocode N = 1000 # 1-in-N packet sampling samples = 40 # sampled packets for one prefix in one minute sampled_bytes = 48000 # sum of their lengths est_packets = samples * N # 40,000 packets est_bytes = sampled_bytes * N # 48,000,000 bytes = 6.4 Mb/s over 60 s rel_se = 1 / sqrt(samples) # about 0.158 ci95 = 1.96 * rel_se # about 0.31: roughly 27,600 to 52,400 packets ``` ## What this means for the operator - Report sampled figures with their sample count or an error bar, never as exact totals. - Cross-check sampled link sums against an exact counter, such as sFlow counter samples or interface counters. - Choose N from the smallest aggregate that must be accurate, not from the link speed alone.
- An sFlow v5 collector loses 20% of datagrams but scales every sample by the configured 1-in-1000. What does it get wrong, and how can it correct it?Every lost datagram removes samples, so the attained rate is lower than configured and the estimates come out about 20% low. Each flow_sample carries sample_pool and a sequence number. From these the collector can compute how many packets the received samples actually represent, and scale by that attained rate instead of the configured one.
- Why can deterministic every-Nth-packet sampling mislead even when N is correct?RFC 5475 warns that systematic sampling can bias results when the traffic has its own periodicity. A stream that sends packets in a regular pattern can line up with the sampler and be consistently over- or under-selected. Random selection gives every packet the same independent chance, which is what the 1/sqrt(c) error estimate assumes.
An exit poll: its margin of error depends on how many voters were actually interviewed, not on what share of the electorate that was. A district where only 40 voters were asked gets a wide margin, however carefully the result is scaled up to the whole district.
saying these in an interview costs you the question
- Multiplying by N turns sampled counts into exact totals.
- Sampling error depends only on N, not on how many samples arrived.
- Every-Nth-packet sampling is always as accurate as random sampling.
- Lost sFlow datagrams do not matter, since the sampling rate is still N.
- Moving from 1-in-1000 to 1-in-10,000 makes small-prefix estimates more precise.