skip to content

Why does an NTP client keep the last eight samples from each server and use the lowest-delay one rather than the newest?

level: middleimportance: should knowfreq 22%

answer

  1. queueing inflates delay
  2. error bounded by half the delay
  3. eight-stage shift register
  4. use each sample only once
  5. dispersion grows with age

basics

~20 s

Queueing delay is what corrupts an NTP offset, and it rarely hits both directions equally. Of the last eight samples, the lowest-delay one was queued least, so its offset has the smallest error bound; the newest may have been queued heavily.

solid answer

~50 s

RFC 5905's **clock filter** runs separately for each server: an 8-stage shift register holds the most recent `(offset, delay, dispersion, time)` tuples, and on each new sample it sorts them by delay and takes the lowest-delay one as that server's current offset. The reason is error, not freshness: the offset error from an asymmetric path is bounded by half the round-trip delay, and extra delay mostly comes from queueing that hits one direction harder, so the least-delayed sample is the least contaminated. Age is handled separately: each stored sample's dispersion grows at `PHI` (15 PPM of elapsed time), and the filter never reuses a sample or accepts one older than the last it used. The spread of the stored offsets gives that server's **jitter**, and the weighted dispersions feed its **synchronisation distance**, which the selection step later uses to compare servers.

go deeper

for a junior

Recall that the client keeps several recent samples from each server and prefers the one that travelled fastest, because slow replies were probably queued.

for a middle

Explain the eight-stage register, sorting by delay, the half-delay error bound, the use-once rule, and how dispersion ages stored samples at 15 PPM.

for a senior

Connect the filter to symptoms: congested links raising jitter, a new client slow to trust a server until about four samples arrive, and spikes the reference code discards.

for a principal

Weigh filtering against responsiveness: a longer sample memory resists queueing noise but reacts later to real change, which shapes poll and burst choices.

## Where the clock filter sits RFC 5905 splits an NTP client's work into a **peer process**, one per server, and a **system process** that looks across all servers. The **clock filter algorithm** (§10) belongs to the peer process. It never compares one server with another; its job is to turn a stream of noisy samples from *one* server into that server's best current estimate. Comparing servers is the later job of the selection, cluster and combine steps. Each sample is a tuple of four values: - **offset** `theta`: how far the server's clock appears to be from the client's; - **delay** `delta`: the round-trip time of that exchange; - **dispersion** `epsilon`: an error bound that grows with time; - **arrival time** `t`. How offset and delay are computed from one request and reply is a separate subject; here they are inputs. ## The eight-stage register The filter is "an 8-stage shift register". Each new tuple is shifted in and the oldest falls out, so the register always holds the eight most recent samples in arrival order. - At start-up every stage holds a **dummy tuple** with dispersion `MAXDISP` (16 s), so an empty filter looks maximally untrustworthy. - If three poll intervals pass with no valid reply, the poll process shifts in a dummy tuple, so a silent server's samples age out rather than staying trusted forever. - Each time the register shifts, the reference code adds `PHI` times the elapsed time to every stored sample's dispersion. `PHI`, the frequency tolerance, is 15 PPM, so a stored sample's error bound grows by 15 µs for every second it waits. ## Choosing the sample On each new sample the filter: 1. Copies the eight stages to a temporary list and **sorts it by increasing delay**. 2. Takes the first entry, the lowest-delay sample, as the server's offset and delay. 3. Exits without changing anything if that sample is **not newer than the last one used**. The reference code calls this the prime directive: use a sample only once, and never one older than the latest used, except before first synchronisation. 4. Computes the server's dispersion and jitter from the list. 5. In RFC 5905's reference code (not its normative text), runs a **popcorn spike suppressor**: if the new offset differs from the previous one by more than `SGATE` (3) times the jitter and little time has passed, the sample is discarded as a spike. ## Why the lowest delay wins A reply delayed by queueing on the way out is usually not delayed equally on the way back, and the offset calculation assumes the two directions are equal. The resulting offset error is bounded by half the round-trip delay, so a sample's delay is a direct measure of how wrong its offset could be. Consider five of the eight stored samples from one server, newest last: | Sample | Delay | Offset | Worst-case error from asymmetry | |---|---|---|---| | 1 | 4.2 ms | +0.9 ms | 2.10 ms | | 2 | 18.6 ms | +6.1 ms | 9.30 ms | | 3 | 4.1 ms | +1.0 ms | 2.05 ms | | 4 | 4.4 ms | +1.2 ms | 2.20 ms | | 5 | 25.0 ms | -8.3 ms | 12.50 ms | Using the newest sample would pull the clock 8.3 ms the wrong way on the strength of a reply that sat in a queue. The filter uses sample 3, with an error bound of about 2 ms; samples 2 and 5 are consistent with it once their much wider bounds are counted. ## What the filter hands on From the sorted list the filter produces the numbers the system process needs: - **Peer dispersion**: the stage dispersions weighted by `1/2^(i+1)`, so the best samples count most. With all dummies it is just under 16 s; each valid sample roughly halves it, and RFC 5905 notes that after the fourth valid sample it is usually a little under 1 s, the `MAXDIST` threshold selection requires. With four near-zero samples, the four remaining dummies contribute 16 × (1/32 + 1/64 + 1/128 + 1/256) = 0.9375 s. - **Jitter**: the RMS difference between the best sample's offset and the others, never below the system precision. - **Synchronisation distance** `lambda = delta/2 + epsilon`, which keeps growing at `PHI` between updates, so a server that stops answering eventually becomes unfit for synchronisation. ## What it does not do - It does not decide which server is right; that is selection. - It does not compute the offset of one exchange; that is the four-timestamp arithmetic. - It does not move the clock; the chosen, combined offset goes to the clock discipline.

  • Why does a freshly started NTP client take several polls before it trusts a server, and what shortens that?
    The filter starts full of dummy tuples with 16 s dispersion, and each valid sample only roughly halves the weighted dispersion. RFC 5905 notes it falls just under the 1 s `MAXDIST` threshold after about four valid samples. RFC 5905 also describes an `IBURST` option that sends a burst of eight packets 2 s apart to a server that is not yet reachable, filling the filter in seconds instead of four poll intervals.
  • What does the popcorn spike suppressor in RFC 5905's reference code protect against?
    Single wild samples. If a new best offset differs from the previous one by more than `SGATE` (3) times the server's jitter, and less than twice the poll interval has passed, the code drops it instead of passing it on to selection. An isolated spike never reaches the clock, while a genuine, lasting change passes once enough time has gone by.

saying these in an interview costs you the question

  • The client always uses the newest reply because it is the freshest.
  • The clock filter averages the offsets of all eight stored samples.
  • The clock filter compares different servers to find the best one.
  • A sample's offset can be wrong by up to its whole round-trip delay.
  • Stored samples stay as trustworthy as new ones until they are shifted out.