skip to content

Why does an NTP client save its learned clock frequency to a file, and what happens at restart when that file is missing?

level: middleimportance: nice to knowfreq 10%

answer

  1. learning a rate takes time
  2. two start states in RFC 5905
  3. NSET versus FSET
  4. a 900 s measurement window

basics

~20 s

Measuring an oscillator's frequency error takes time. With a saved frequency, RFC 5905's discipline starts in FSET and disciplines at once; without one it starts in NSET and spends the 900 s stepout interval measuring frequency first.

solid answer

~50 s

The discipline's most expensive product is the **frequency correction**: how many ppm the local oscillator is off. Implementations save it to a frequency file, often called the drift file, so a restarted client does not have to learn it again. RFC 5905's skeleton starts in `FSET` when that value exists and in `NSET` when it does not. From `NSET`, the first update moves the client to `FREQ`, where it waits out the stepout interval `WATCH` (900 s) and computes the frequency directly as the change in offset over that interval. From `FSET`, it goes straight to normal operation with the saved frequency. In both states an initial offset above the 125 ms step threshold is stepped at the first update; what the file buys is about 15 minutes of disciplined frequency, and a saved value learned on different hardware has to be unlearned.

go deeper

for a junior

Remember that the file holds the oscillator's learned rate error, not the time, and that it saves the client from relearning it after a restart.

for a middle

Walk through NSET versus FSET, the 900 s frequency measurement, and compute a frequency error from two offsets taken an interval apart.

for a senior

Diagnose clients that step after every start or migration by asking whose oscillator the saved frequency describes, and decide when to remove it.

for a principal

Consider how image-based fleets, host moves and hardware swaps break the assumption that a saved frequency belongs to one machine, and how to manage that at scale.

## What the discipline has to learn An NTP client's **clock discipline** (RFC 5905 section 11.3) corrects two things: the **phase**, how far the clock is off now, and the **frequency**, how fast that error grows because the local oscillator is not exactly on its nominal rate. Phase is cheap: one good sample tells you the offset. Frequency is expensive: it is the slope of offset over time, so it needs samples spread over a long interval. RFC 5905 says that a purely linear feedback loop starting without knowledge of the intrinsic frequency takes **several hours** to develop an accurate measurement, and its non-linear start-up state machine does it in **15 minutes**. Once learned, that frequency is worth keeping. Implementations write it to a file (commonly called the **drift file**; RFC 5905's skeleton just says "frequency file") so that the next start can use it. ## Two ways to start RFC 5905's appendix skeleton initialises the clock with one test: if a frequency file exists, load the frequency and enter `FSET`; otherwise enter `NSET`. | Start state | Meaning | First update with offset below 125 ms | First update with offset above 125 ms | |---|---|---|---| | `NSET` | clock never set, no frequency | record the offset, enter `FREQ` | step the time, enter `FREQ` | | `FSET` | frequency loaded from file | enter `SYNC`, adjust phase | step the time, enter `SYNC` | From `FREQ`, the client ignores updates until the stepout threshold `WATCH` (900 s) has passed, then computes the frequency directly and moves to `SYNC`: 1. At the first update the offset is, say, 2 ms. 2. 900 s later it is 47 ms, because the uncorrected oscillator has kept drifting. 3. The frequency error is (47 - 2) ms / 900 s = 45 ms / 900 s = 5 x 10^-5 = **50 ppm**. 4. The client applies that correction and enters normal operation. With a frequency file the client skips all four steps and starts at step 4. ## What the file buys, and what it does not - **It does not decide whether the clock is stepped at start.** In both `NSET` and `FSET` an initial offset above the step threshold is stepped at the first update; the skeleton's comment notes that operators get nervous if setting the clock the first time takes 17 minutes. - **It does not hold the time of day.** It holds a rate. The time still has to come from the network after every restart. - **It buys disciplined frequency from the first update**, instead of a quarter of an hour during which the clock runs at its raw oscillator error. - **It can let the poll interval grow sooner**, because a client with a good frequency sees small, stable offsets. ## When the saved value is wrong The file describes one oscillator. If the hardware changes, it describes the wrong one: - a board or machine replaced, with the old disk kept; - a virtual machine moved to a different physical host, whose oscillator it now runs on; - a file copied from a template image to many machines. The client starts in `FSET` and trusts the stale value. It still converges, through the ordinary feedback loop, but meanwhile offsets grow between polls and may cross the step threshold, producing a spike, a wait, and a step. RFC 5905's skeleton clamps any frequency correction to plus or minus 500 ppm (`MAXFREQ`), so a badly wrong file is bounded, but it can still be worse than starting fresh. Removing a file known to be stale forces the 15-minute direct measurement of `NSET`, which is often the faster route back. ## Operating it - Keep the file on storage that survives restarts and is specific to the machine. - Do not bake it into images shared by many machines. - After a hardware or host change, expect a period of re-learning, or remove the file. - When an NTP client steps shortly after every start, check whether its saved frequency belongs to this machine. ## Common mistakes - Thinking the drift file stores the last known time. - Thinking a client cannot synchronise without it. - Treating it as portable configuration.

  • Can a saved frequency file ever be worse than having none?
    Yes. If it was learned on different hardware, its error can exceed the new oscillator's own error. The client starts in `FSET`, trusts it, and has to unlearn it through the ordinary feedback loop, which RFC 5905 notes can need hours to measure a frequency, while offsets build up and may trigger a spike and a step. Deleting a file known to be stale forces `NSET`'s 15-minute direct measurement.
  • Does a frequency file change whether an NTP client steps its clock at start-up?
    No. In both `NSET` and `FSET`, RFC 5905's skeleton steps an initial offset above the 125 ms step threshold at the first update. The file changes what follows: `FSET` goes straight to normal discipline, while `NSET` spends the 900 s stepout interval measuring the oscillator's frequency first.

saying these in an interview costs you the question

  • The drift file stores the last known time so the clock resumes after reboot
  • Without a frequency file an NTP client cannot synchronise at all
  • A frequency file can be copied between machines like any configuration file
  • The frequency file decides whether the clock is stepped at start-up