skip to content

Clock Drift and Discipline

Oscillators drift by parts per million, so the daemon learns the error and slews rather than steps, keeping time from jumping back. Leap-second smearing and step-at-boot are the usual probes.

on this pageshow

questions

5

Why does a computer clock left without NTP drift, and how far does a 50 PPM oscillator error move it in a day?

level: juniorimportance: must knowfreq 38%

answer

  1. no crystal runs exactly on frequency
  2. parts per million of seconds per second
  3. 86,400 seconds in a day
  4. correct the rate, not only the time

basics

~20 s

A clock counts oscillator ticks, and no oscillator runs at exactly its nominal frequency, so the error accumulates. 50 PPM is 50 microseconds per second: 50 x 10^-6 x 86,400 s is about 4.3 seconds a day.

solid answer

~50 s

A computer clock is an oscillator plus a counter, and the oscillator's real frequency is always slightly off its nominal value. RFC 5905 measures that error in parts per million, where 1 ppm is 10^-6 seconds per second, so 1 ppm costs about 86.4 ms a day and 50 ppm costs 50 x 86.4 ms, about **4.32 s a day** or half a minute a week. The error also moves with temperature and age, so setting the clock once does not hold. That is why an NTP client disciplines two things: the **offset** (how wrong the time is now) and the **frequency** (how fast it is going wrong), so the clock stays close between polls. When sources vanish, RFC 5905 assumes the clock may wander at up to `PHI` = 15 ppm and grows its advertised dispersion by about 1.3 s a day.

go deeper

for a junior

Recall that clocks drift because oscillators are never exactly on frequency, and be able to turn ppm into seconds per day by multiplying by 86,400.

for a middle

Explain that NTP's discipline corrects both offset and frequency, and why a learned frequency keeps the clock close between polls and lets polling slow down.

for a senior

Reason about holdover: what a server's clock does when its sources vanish, how temperature changes the rate, and how dispersion growth at 15 ppm signals declining trust downstream.

for a principal

Weigh when commodity oscillators plus NTP are enough and when a requirement for tighter holdover justifies better oscillators or a different time distribution design.

## Why a clock drifts at all Every computer clock is built the same way: an **oscillator** (usually a quartz crystal) ticks at a nominal frequency, and a counter turns ticks into seconds. If the crystal ran at exactly its nominal frequency, the clock would keep perfect time once set. It never does. Manufacturing tolerance puts every crystal slightly off its label, and the error is not even constant: - **Temperature** changes the frequency; RFC 5905 notes that a temperature spike can cause a frequency surge the discipline has to chase. - **Aging** slowly shifts it over months and years. - **Load and power state** change the temperature inside the machine, and so the rate. RFC 5905 (NTPv4, which obsoletes RFC 1305 and the SNTP document RFC 4330) models the time error as `T(t) = T(t0) + R(t0)(t - t0) + 1/2 D(t0)(t - t0)^2 + e`, where `R` is the **frequency offset** and `D` the **aging rate**. For computer oscillators the aging term is ordinarily neglected, which leaves the useful picture: time error grows linearly with the frequency error. ## Parts per million, worked RFC 5905 keeps frequencies in seconds per second and expresses them in **parts per million (ppm)**, where 1 ppm = 10^-6 s/s. A clock 1 ppm fast gains one microsecond every second. Multiply by the 86,400 seconds in a day: | Frequency error | Per day | Per week | |---|---|---| | 1 ppm | 86.4 ms | 0.60 s | | 15 ppm | 1.30 s | 9.07 s | | 50 ppm | 4.32 s | 30.24 s | | 100 ppm | 8.64 s | 60.48 s | | 500 ppm | 43.2 s | 302.4 s | So a 50 ppm oscillator, which looks like a tiny number, moves the clock about **4.3 seconds a day**. Two servers drifting in opposite directions at that rate disagree by nearly nine seconds after one day, which is enough to reorder log lines and confuse anything that compares timestamps across machines. ## What NTP does about it: phase and frequency An NTP client does not merely copy the server's time. Its **clock discipline** algorithm (RFC 5905 section 11.3) is a feedback loop that corrects two quantities: 1. **Phase (offset)** - how far the clock is from the server's time right now. 2. **Frequency** - how fast that offset is growing. Once the frequency correction is learned, the clock runs at nearly the right rate on its own, and the offset between polls stays small. That is what lets an NTP client stretch its poll interval (the poll algorithm itself belongs to source selection and polling): a clock that barely drifts needs fewer samples. RFC 5905's appendix skeleton clamps the frequency correction it will apply to plus or minus 500 ppm (its `MAXFREQ` constant), so an oscillator worse than that cannot be fully disciplined by that design. ## What happens when sources disappear If every server becomes unreachable, a disciplined clock keeps running at its last learned frequency. It now drifts only by the residual error plus whatever temperature and aging change afterwards, which is usually far less than the raw crystal error. RFC 5905 cannot know that residual, so it assumes a worst case: the **frequency tolerance `PHI` = 15 ppm**. The client's dispersion, part of the error bound it reports to its own clients, grows at `PHI`, which RFC 5905 works out as about **1.3 s per day**. Downstream clients see the server's time becoming less trustworthy and can prefer another source. ## Typical accuracy once disciplined RFC 5905 gives indicative figures: primary servers on modern machines are precise within a few tens of microseconds, secondary servers and clients on fast LANs within a few hundred microseconds at poll intervals up to 1024 seconds, and within a few tens of milliseconds with poll intervals stretched up to 36 hours. Over long network paths the limit is usually the path rather than the oscillator: the error an asymmetric path adds is a property of the offset calculation, not of discipline. ## Common mistakes - Reading 50 ppm as 50 microseconds in total rather than 50 microseconds every second. - Assuming a new machine keeps good time and drift comes only with age. - Believing one correction at boot holds forever; the rate changes with temperature. - Treating NTP as offset-only; without a learned frequency the clock drifts away again between polls.

  • If a disciplined server loses all its NTP sources, how quickly does its clock go wrong?
    It keeps running at its last learned frequency, so it drifts only by the residual error plus later temperature and aging changes, often far less than the raw crystal error. RFC 5905 cannot know that residual, so it assumes the worst-case `PHI` of 15 ppm and grows the dispersion it advertises by about 1.3 s a day, warning downstream clients that its time is losing trust.
  • Why not simply poll the server every second instead of learning the frequency?
    Every sample carries network jitter, so correcting the phase on each noisy sample would chase the noise and load the server. Learning the frequency lets the client average over longer intervals and poll less often. Polling faster also does nothing about the error an asymmetric network path adds; that comes from the path, not from the poll rate.

saying these in an interview costs you the question

  • A new server's clock is accurate; drift only starts as the hardware ages
  • 50 PPM means the clock ends up 50 microseconds off in total
  • Setting the clock once at boot is enough for a server's lifetime
  • NTP only corrects the time offset and leaves the clock's rate alone
  • The oscillator's frequency error is fixed, so one measurement corrects it forever
open as a page

When an NTP client measures its clock offset, when does it slew and when does it step the clock, and why does that matter?

level: middleimportance: must knowfreq 32%

basics

~20 s

Below RFC 5905's 125 ms step threshold an NTP client slews, running the clock slightly fast or slow so time never jumps. In normal operation larger offsets are stepped only after persisting 900 s; beyond 1000 s it should exit.

open as a page

A cluster's NTP servers must cross a leap second; how does inserting the second differ from smearing it, and why must clients never mix the two?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Inserting adds a 61st second to the month's final minute, announced by NTP's leap indicator. Smearing runs servers slightly slow for hours, so no second repeats. Mixed sources disagree by up to a second, which RFC 8633 forbids.

open as a page

When an NTP client in a VM resumes seconds off after live migration, how does discipline correct it, and when should stepping be allowed?

level: seniorimportance: should knowfreq 15%

basics

~20 s

A multi-second offset exceeds RFC 5905's 125 ms step threshold, so the client treats it as a spike and steps after 900 s; past 1000 s it should exit. A common policy steps at boot and alerts on later steps.

open as a page

Why does an NTP client save its learned clock frequency to a file, and what happens at restart when that file is missing?

level: middleimportance: nice to knowfreq 10%

basics

~20 s

Measuring an oscillator's frequency error takes time. With a saved frequency, RFC 5905's discipline starts in FSET and disciplines at once; without one it starts in NSET and spends the 900 s stepout interval measuring frequency first.

open as a page