skip to content

After NTP's selection step leaves five truechimers a few milliseconds apart, how do the cluster and combine steps produce one offset and pick the system peer?

level: seniorimportance: should knowfreq 14%

answer

  1. trim the outlier, round by round
  2. selection jitter against peer jitter
  3. stop at three survivors
  4. stratum first, then distance
  5. weight by inverse root distance

basics

~20 s

NTP's cluster step repeatedly discards the survivor whose offset sits furthest from the others, until that stops helping or three remain; the combine step averages the rest weighted by inverse root distance, and the top-ranked survivor becomes the system peer.

solid answer

~50 s

RFC 5905's **cluster** algorithm ranks survivors by a merit factor, `stratum × MAXDIST + root distance`; `MAXDIST` is 1 s, so stratum dominates. It then runs rounds: compute each survivor's **selection jitter**, the RMS of its offset against the others, and discard the worst, stopping when that worst value is below the smallest peer jitter, since discarding more cannot help, or when only `NMIN` (3) remain. The first survivor in rank order becomes the **system peer**: the client inherits its stratum plus one, its reference ID and its root values. The **combine** algorithm then averages the survivors' offsets weighted by the reciprocal of each one's root distance, so a close, low-dispersion server counts more than a distant one even if it ranks lower. RFC 5905's reference code also avoids clock hopping by keeping the old system peer when it ties on stratum with the new first survivor.

go deeper

for a junior

Recall that after wrong servers are thrown out, NTP still trims outliers and averages several good servers rather than trusting one.

for a middle

Explain selection jitter, the two stopping rules, the stratum-dominated merit factor, and inverse-distance weighting in the combine step.

for a senior

Use these steps to explain field behaviour: why a nearby higher-stratum server never becomes system peer, why the reference ID changes, and how hopping between peers adds jitter.

for a principal

Weigh combining many sources against tracking one excellent source: averaging cuts independent noise, but mixing sources of very different quality can dilute the best one.

## Where these steps sit An RFC 5905 client mitigates among its servers in three steps in the **system process**. **Selection** has already discarded falsetickers, servers whose correctness intervals lie outside the majority. What remains are truechimers, all plausibly correct, but not equally good: some are noisier, further away, or slightly off from the rest. The **cluster** step (§11.2.2) thins them statistically, and the **combine** step (§11.2.3) turns the survivors into one offset for the clock discipline. Three quantities appear throughout: - **Offset** `theta_p`: the server's current offset from its clock filter. - **Peer jitter** `psi_p`: how much that server's own recent samples scatter. - **Root distance** `lambda`: the maximum error from all causes between the client and the primary reference behind that server. ## Ranking: the merit factor Each survivor gets a merit factor `stratum × MAXDIST + lambda`, with `MAXDIST` equal to 1 s, and the list is sorted by it, lowest first. Because one stratum step is worth a full second of distance, and real root distances are milliseconds, **stratum decides the order** and distance only breaks ties within a stratum. The reference code says it plainly: the metric "is dominated first by stratum and then by root distance". ## Clustering, round by round 1. For each survivor compute its **selection jitter** `psi_s`: the RMS of the differences between its offset and every other survivor's. 2. Find the survivor with the largest `psi_s`, the statistical outlier. 3. Stop if that largest `psi_s` is **smaller than the smallest peer jitter** of any survivor, because the outlier is then within the noise every server already has and discarding it gains nothing. 4. Also stop if only `NMIN` (3) survivors remain. 5. Otherwise discard the outlier and go back to step 1. The final largest selection jitter is kept as the system selection jitter `PSI_s`. ## A worked example Five truechimers, offsets and root distances in milliseconds: | Server | Stratum | Offset | Root distance | Peer jitter | Merit (s) | |---|---|---|---|---|---| | A | 1 | +1.0 | 10 | 0.5 | 1.010 | | D | 2 | +4.6 | 15 | 0.4 | 2.015 | | B | 2 | +1.4 | 20 | 0.5 | 2.020 | | C | 2 | +0.7 | 40 | 0.6 | 2.040 | | E | 3 | +1.2 | 8 | 0.5 | 3.008 | - **Round 1:** D's selection jitter is about 3.53 ms, the largest and above the smallest peer jitter (0.4 ms). D is discarded. - **Round 2:** C's is about 0.53 ms, still above the smallest remaining peer jitter (0.5 ms). C is discarded. - **Round 3:** three survivors remain, equal to `NMIN`, so clustering stops with A, B and E. ## Choosing the system peer The first survivor in merit order, **A**, becomes the system peer. Its variables become the client's system variables: stratum is A's plus one, the reference ID and reference time are A's, and root delay and dispersion build on A's. This is what a downstream client sees when it queries this one. RFC 5905's reference code adds one refinement: if the previous system peer has the same stratum as the new first survivor, it is kept, to avoid **clock hopping** between near-equal servers. ## Combining The combine algorithm weights each survivor's offset by `1 / lambda` and normalises: - weights: A 1/10, B 1/20, E 1/8, which is 0.100, 0.050 and 0.125, summing to 0.275; - weighted offsets: 0.100 × 1.0 + 0.050 × 1.4 + 0.125 × 1.2 = 0.320; - combined offset: 0.320 / 0.275 ≈ **+1.16 ms**. E ranked last but carries the largest weight, about 45%, because its root distance is shortest. Rank chooses whose identity the client inherits; distance decides whose offset counts most. The system jitter is then `PSI = sqrt(PSI_s^2 + PSI_p^2)`, combining the selection jitter with the peer jitter. ## Why not trust only the best server - Averaging several truechimers reduces independent noise that any single server's path adds. - Keeping at least three survivors means losing one does not leave a single unchecked source. - The reference code notes the trade-off: jitter can sometimes be lower without combining, "especially when frequent clockhopping is involved", and lets an operator designate a preferred peer instead.

  • Why does the cluster step stop at three survivors instead of keeping only the single best server?
    `NMIN` (3) keeps enough survivors for the combine step to average out independent path noise, and losing one still leaves more than a single unchecked source. Clustering also stops earlier if the worst selection jitter is already below the smallest peer jitter, because removing that survivor would no longer improve the combined offset.
  • Why does a nearby stratum-3 server never become system peer while stratum-2 survivors exist, and does it matter?
    The merit factor is `stratum × MAXDIST + root distance`, with `MAXDIST` 1 s, so each stratum step outweighs up to a second of distance and the lower stratum ranks first. It matters less than it seems: the combine step weights offsets by inverse root distance, so the close server still contributes most to the final offset; it just does not supply the stratum and reference ID.

saying these in an interview costs you the question

  • NTP picks the single lowest-delay server and ignores the others.
  • The combine step gives every surviving server an equal vote.
  • The cluster step discards the highest-stratum server first.
  • Clustering keeps discarding until only the best server is left.
  • The system peer is always the server with the smallest root distance.