skip to content

Source Selection and Polling

Adaptive polling, a filter that keeps the best recent samples, and an intersection step that outvotes a lying server. 'Why configure four time servers, not two?' is the question it answers.

on this pageshow

questions

6

When one of an NTP client's two time servers starts serving time 3 s off, what does the client do, and why recommend four servers?

level: middleimportance: must knowfreq 35%

answer

  1. an interval around each offset
  2. a majority must overlap
  3. falsetickers fewer than half
  4. one down and one lying

basics

~20 s

Two NTP servers that disagree cannot outvote each other, so the client finds no majority and stops correcting its clock from either. Four independent sources let a client still outvote one wrong server after another has become unreachable.

solid answer

~50 s

RFC 5905's selection step gives each server a **correctness interval**, its offset plus or minus its root synchronisation distance, and looks for an interval shared by a majority. It first assumes no falsetickers, then one more each round, and gives up once the assumed number `f` is no longer less than half the candidates. Two servers 3 s apart have intervals tens of milliseconds wide that cannot overlap, and `f = 1` is not less than `2/2`, so no majority exists and the client stops disciplining its clock; it free-runs while an operator works out which server is wrong. With three servers, the two good ones outvote the bad one. RFC 8633 says operators SHOULD use at least four independent, diverse sources: with four, one can be unreachable and the remaining three still outvote a single wrong server.

go deeper

for a junior

Remember the rule of thumb: two clocks that disagree cannot tell you which is right; a third breaks the tie and a fourth covers one being offline.

for a middle

Walk through the intersection step: correctness intervals, assumed falsetickers starting at zero, the rule that they must be fewer than half, and apply it to two, three and four servers.

for a senior

Count failures as production produces them: one source unreachable while another is wrong, servers sharing an upstream, and smeared and unsmeared servers forming two equal cliques.

for a principal

Treat source count as a fault budget: decide how many simultaneous unreachable and wrong sources the estate must survive, then buy independent sources to cover it.

## The scenario A fleet of hosts is configured with two internal NTP servers. One of them loses its reference and starts serving time that is **3 s off**, while still answering every request promptly. Each client now holds two confident, contradictory answers. What it does next is decided by RFC 5905's **selection algorithm** (§11.2.1), and the outcome is the clearest argument for configuring more servers. ## Correctness intervals For every usable server the client has an **offset** `theta` (how far the server's clock is from its own, taken from that server's best recent sample) and a **root synchronisation distance** `lambda`, the maximum error from all causes between the client and the primary reference behind that server. The pair defines a correctness interval: - lowpoint `theta - lambda` - midpoint `theta` - highpoint `theta + lambda` The true time should lie inside the interval of every correct server, so the intervals of truechimers overlap. On a healthy LAN `lambda` is typically milliseconds to tens of milliseconds, so two servers 3 s apart produce intervals with a wide gap between them. ## The intersection step RFC 5905 describes a refinement of Marzullo's agreement algorithm: 1. Place the three points of every candidate's interval on a list and sort it. 2. Assume `f` falsetickers, starting at `f = 0`. 3. Scan upward for the lowest point that at least `m - f` intervals contain, where `m` is the number of candidates; scan downward for the highest such point. 4. If an interval `[l, u]` is found with `l < u` and no more midpoints outside it than the assumed falsetickers, it is the **majority clique**; candidates inside it are the truechimers. 5. Otherwise add one to `f`. If `f` is still less than `m / 2`, try again; if not, selection **fails**. When selection fails, RFC 5905 §11.2 says the system process "exits without disciplining the system clock". ## Counting servers Applying step 5's rule, `f < m / 2`, to common configurations: | Reachable servers | Wrong servers | Outcome | |---|---|---| | 2 | 1 | `f = 1` is not `< 1`: no majority, no discipline | | 3 | 1 | `f = 1 < 1.5`: the two good servers outvote it | | 3 of 4 (one unreachable) | 1 | same as three: outvoted | | 2 of 3 (one unreachable) | 1 | same as two: no majority | | 4 | 2 that agree with each other | `f = 2` is not `< 2`: two cliques of two, no majority | | 5 | 2 | `f = 2 < 2.5`: outvoted | An unreachable server is not a candidate at all, so `m` shrinks when one goes silent. That is the core of RFC 8633's advice: "Four sources will provide sufficient backup in case one source goes down", and with three left the client can still outvote one wrong one. Operators "SHOULD use at least four independent, diverse sources of time". ## What the operator sees - With two servers, clients stop being disciplined at the moment the second server disagrees. Their clocks drift on the local oscillator until a human decides which server is wrong. - With four, the clients quietly mark the bad server a falseticker and keep good time; the fault shows up only in monitoring. RFC 8633 §3.5 says operators SHOULD monitor their time sources, because the algorithm hides the failure from the clients. - Adding servers that share one upstream reference does not add independent votes. ## Majority versus the Byzantine bound The selection step needs a strict majority, so tolerating `f` wrong servers needs at least `2f + 1` candidates. Classic Byzantine agreement needs `3f + 1`, and RFC 5905's reference code comments that "ordinarily, the Byzantine criteria require four survivors". For a single fault, the majority rule alone asks for three; adding a spare for an unreachable server, as RFC 8633 does, or applying the Byzantine bound both arrive at four. The algorithm also cannot help when the wrong servers are the majority: RFC 8633's analysis "assumes that a majority of the servers used in the solution are honest". ## A real failure of the two-against-two kind RFC 8633 §3.2 records that around the June 2015 leap second some operators leap-smeared their servers and others did not. Clients with two of each saw two consistent pairs and "could not determine an accurate time source": the four-server row with two agreeing wrong servers, produced by mixing incompatible time sources rather than by a fault.

  • Why can't the client break the tie by preferring the server with the lower stratum or the lower delay?
    Stratum says how many hops a server is from a reference, and delay says how noisy the path is; neither says whether its time is right. A misconfigured stratum-1 server can be confidently wrong. RFC 5905 uses stratum and distance only to rank servers after selection has removed falsetickers, so with two disagreeing servers there is nothing left to rank.
  • Why did some clients with four servers lose sync around the June 2015 leap second?
    RFC 8633 §3.2 describes it: two of their four servers leap-smeared and two did not. Each pair agreed internally, giving two cliques of two. Selection needs `f < m / 2`; with `m = 4`, `f = 2` fails, so no majority existed. The lesson is not to mix smeared and unsmeared sources behind one client.

saying these in an interview costs you the question

  • When two servers disagree, the client averages them and halves the error.
  • RFC 5905 keeps following the current system peer when its servers disagree.
  • Three servers tolerate one unreachable and one wrong server at the same time.
  • Four servers are recommended so that two wrong servers can be outvoted.
  • The lower-stratum server always wins a disagreement between two servers.
open as a page

Why do operators give an NTP client several time servers instead of one, beyond simply having a spare if one fails?

level: juniorimportance: should knowfreq 42%

basics

~20 s

With one server, an NTP client copies whatever error that server has and cannot notice it. With several, the client compares them, discards a server whose time disagrees with the majority (a falseticker) and averages the rest.

open as a page

How does an NTP client choose its poll interval between MINPOLL and MAXPOLL, and why does it lengthen the interval once its clock is stable?

level: middleimportance: should knowfreq 25%

basics

~20 s

An NTP poll interval is a power of two seconds. RFC 5905 raises the exponent while measured offsets stay small relative to jitter and lowers it when they grow, so a stable clock polls rarely, averaging over longer spans and loading servers less.

open as a page

Why does an NTP client keep the last eight samples from each server and use the lowest-delay one rather than the newest?

level: middleimportance: should knowfreq 22%

basics

~20 s

Queueing delay is what corrupts an NTP offset, and it rarely hits both directions equally. Of the last eight samples, the lowest-delay one was queued least, so its offset has the smallest error bound; the newest may have been queued heavily.

open as a page

After NTP's selection step leaves five truechimers a few milliseconds apart, how do the cluster and combine steps produce one offset and pick the system peer?

level: seniorimportance: should knowfreq 14%

basics

~20 s

NTP's cluster step repeatedly discards the survivor whose offset sits furthest from the others, until that stops helping or three remain; the combine step averages the rest weighted by inverse root distance, and the top-ranked survivor becomes the system peer.

open as a page

Why can NTP's majority-based source selection not protect a client whose servers are mostly attacker-controlled, and what does Khronos (RFC 9523) change?

level: seniorimportance: nice to knowfreq 6%

basics

~20 s

NTP's selection outvotes a minority of bad servers but follows a majority, so an attacker controlling most of a client's few servers can shift its time. Khronos, an Informational watchdog, samples random servers from a large pool and trims outliers.

open as a page