Why do operators give an NTP client several time servers instead of one, beyond simply having a spare if one fails?
answer
- a clock cannot check itself
- one source passes its error through
- compare, discard, combine
- falseticker versus truechimer
basics
~20 sWith one server, an NTP client copies whatever error that server has and cannot notice it. With several, the client compares them, discards a server whose time disagrees with the majority (a falseticker) and averages the rest.
solid answer
~50 sA client given a single server has nothing to judge it against: RFC 8633 says that any issue with the time at that one source is passed on to the client. Given several servers, an RFC 5905 client runs its **mitigation algorithms**: a selection step looks for a majority of servers whose error intervals overlap and casts out any server outside it as a **falseticker**; a cluster step trims statistical outliers; a combine step averages the survivors, weighting the more trustworthy ones more. So extra servers buy three things: redundancy when one becomes unreachable, *detection* of a server that answers promptly but with wrong time, and a less noisy estimate. The protection needs enough independent servers to form a majority, which is why RFC 8633 recommends at least four, and it needs a full NTP client: RFC 5905 does not require an SNTP client to implement these algorithms.
go deeper
Recall that one server's error passes straight to the client, and that several servers let the client spot and drop one that disagrees, not only survive an outage.
Explain the stages: a per-server filter, selection of a majority that agrees, clustering out outliers, then a weighted average, and why each needs more than one source.
Show that you check source independence: servers sharing one upstream, firmware or path can be wrong together, and a majority vote only protects against independent failures.
Frame source count and diversity as a risk budget: which failures the estate must survive, what each extra independent source costs, and where majority voting stops helping.
## What one source gives you An NTP client learns the time by exchanging packets with a server and estimating how far its own clock is from the server's. That estimate is only as good as the server. If the server's reference has failed, its clock was set wrong, or it is misconfigured, it still answers on time and with a plausible-looking timestamp. A client that knows only that server has **no second opinion**, so it follows the server wherever it goes. RFC 8633, the NTP Best Current Practice, walks through this directly: with one source "any issue with the time at the source will be passed on to the client". With two sources that disagree, "it will be difficult to know which one is correct without making use of information from outside of the protocol". The value of several servers is therefore not only availability; it is the ability to **notice** that one of them is wrong. ## Truechimers and falsetickers RFC 5905 uses two terms for this: - A **truechimer** is a clock that keeps time consistent with a trusted standard. - A **falseticker** is a clock that shows misleading or inconsistent time, even though it may answer every request promptly. The client cannot see UTC directly; it only sees the servers. So it decides who is a falseticker by **agreement**: servers whose estimates overlap form a majority, and a server that sits outside that majority is cast out. This is why the count matters: a falseticker can only be outvoted if there are enough truechimers to outvote it. ## What a full NTP client does with several sources An RFC 5905 client processes its servers in stages: 1. **Clock filter, per server.** It keeps the last eight samples from each server and uses the lowest-delay one, which gives each server a current offset and an error estimate. 2. **Selection.** It builds a correctness interval around each server's offset and searches for an interval shared by a majority; servers outside it are falsetickers. 3. **Cluster.** Among the survivors, it repeatedly discards the one whose offset sits furthest from the rest, while that still reduces the error. 4. **Combine.** It averages the remaining offsets, weighting each by the reciprocal of its root synchronisation distance, and passes that single value on to the clock discipline. None of steps 2-4 does useful work with one server: step 2 succeeds trivially, because a lone interval always overlaps itself, and there is nothing to cluster or average. ## One source versus several | Property | One server | Several independent servers | |---|---|---| | Survives an outage | No; the client free-runs | Yes, while enough remain reachable | | Detects a wrong server | No; the error passes through | Yes, if the wrong one is a minority | | Noise in the final offset | That server's noise alone | Reduced by weighted averaging | | Who decides which is right | Nobody | The majority, by overlapping intervals | ## Limits worth knowing - **Majority, not magic.** RFC 8633 notes the analysis "assumes that a majority of the servers used in the solution are honest". If most of the configured servers are wrong in the same way, the client follows them. - **Independence.** RFC 8633 §3.3 recommends a diversity of reference clocks. Servers that share one upstream source, one firmware version or one chipset can fail together and agree on the same wrong time. - **Default acceptance of one source.** RFC 5905 sets `CMIN`, the minimum number of candidates, to one "for historic reasons", noting that suspicious operators would raise it. A client given a single server therefore synchronises to it rather than refusing. - **Full NTP only.** RFC 5905 §14 says SNTP clients "do not need to implement the mitigation algorithms". An SNTP client is intended for a single upstream server, so listing several servers does not give it majority voting. ## The practical answer Configure several servers because time can be **wrong** without being **down**. Redundancy is the obvious benefit; detection and averaging are the reasons NTP's design asks for a set of sources, and they only work with enough independent servers that a single bad one is a minority. How many that is, and why the answer is four rather than two, follows from the selection step's majority rule.
- If an RFC 5905 client is given one server that is 3 s wrong, will it refuse to use it?Not by default. With one candidate the selection step finds an intersection trivially, because the server's interval overlaps itself, and RFC 5905 sets the minimum candidate count, `CMIN`, to one for historic reasons. The client follows that server. How it then moves its own clock by 3 s, stepping or slewing, is the clock discipline's business, not source selection's.
- Does adding four servers help if they all take time from the same upstream source?Much less than it seems. Majority voting only catches failures that are independent. If all four servers inherit one upstream reference, firmware or path, one fault makes all four agree on the same wrong time, and the client sees a healthy majority. RFC 8633 §3.3 recommends diverse reference clocks and independent implementations for exactly this reason.
Checking the time with passers-by: ask one person and you take whatever their watch says, right or wrong. Ask four and the one whose watch is twenty minutes off stands out while the other three roughly agree and can be averaged. Ask only two who disagree and you cannot tell whose watch is wrong.
saying these in an interview costs you the question
- Extra NTP servers are only there as a backup in case one goes down.
- A low stratum number proves that a server's time is correct.
- The client simply averages all configured servers, including a wrong one.
- Two servers are enough for the client to vote out a bad one.
- Any client that speaks the NTP wire format runs the selection algorithm.