skip to content

When an NTP client in a VM resumes seconds off after live migration, how does discipline correct it, and when should stepping be allowed?

level: seniorimportance: should knowfreq 15%

answer

  1. a pause the guest never saw
  2. the offset against 125 ms
  3. a spike, then a step after 900 s
  4. panic thresholds and restart loops

basics

~20 s

A multi-second offset exceeds RFC 5905's 125 ms step threshold, so the client treats it as a spike and steps after 900 s; past 1000 s it should exit. A common policy steps at boot and alerts on later steps.

solid answer

~50 s

Live migration pauses the guest; if its clock does not advance during the pause, it resumes behind by roughly the pause. To the NTP client that is an offset above the 125 ms step threshold, so RFC 5905's discipline ignores it as a possible spike and steps only once outliers persist past `WATCH` (900 s), leaving up to 15 minutes of wrong time. The guest also now runs on a different oscillator, so its learned frequency is stale and must be relearned. An offset above `PANICT` (1000 s), as after a long suspension, SHOULD make the client exit. The common policy is to **step at boot**, before applications start, and treat later steps as events to alert on. RFC 8633 warns against ignoring the panic threshold on every restart: with automatic restarts, a forged large offset is then accepted.

go deeper

for a junior

Know that a paused or migrated virtual machine can resume with a wrong clock, and that NTP corrects it by slewing or stepping.

for a middle

Explain which offsets are slewed, which are stepped after the 900 s stepout, and why an offset above 1000 s makes the client stop.

for a senior

Set a fleet policy: step at boot, alert on later steps, keep the panic threshold, and fix migration jumps at the platform rather than by loosening NTP.

for a principal

Balance continuity against correctness across workloads, and decide where clock correction for migrations belongs in the platform's design.

## What migration does to a guest clock During **live migration** a virtual machine is copied to another physical host and paused for a final transfer. How the guest's clock behaves across that pause depends on the virtualisation platform, not on NTP: - If the guest's clock does not advance while paused, it resumes **behind** by roughly the pause length. - After a longer suspend or snapshot restore the guest can be minutes or hours behind. - The guest now runs on the destination host's oscillator, so the **frequency correction** it learned describes the wrong hardware. So a migration hands the NTP client two problems at once: a **phase** error, the immediate offset, and a **frequency** error, a drift rate it has not learned yet. ## How RFC 5905 handles the offset RFC 5905's clock discipline sorts an offset by size: | Offset after resume | What the discipline does | |---|---| | below 125 ms (`STEPT`) | slews it out gradually | | 125 ms to 1000 s | ignores it as a possible spike; steps if outliers persist past 900 s (`WATCH`) since the last accepted update | | above 1000 s (`PANICT`) | SHOULD exit with a diagnostic message; an operator sets the clock | For a guest resuming 8 s behind, the client in normal operation moves to its spike state, ignores the next updates, and steps once 900 s have passed since its last accepted update, so up to about 15 minutes later. After the step RFC 5905 requires that all associations be reset, and the appendix skeleton restarts polling from the minimum interval. For those 15 minutes, every log line and timestamp from that guest is 8 s wrong. ## The frequency side Even after the phase is fixed, the stale frequency keeps biting. Each poll interval the clock drifts by the difference between the two hosts' oscillators; at 30 ppm of difference that is 1.9 ms every 64 s, small enough to slew, but over a long poll interval the error builds. The discipline relearns the frequency through its feedback loop, and RFC 5905 notes that its linear loop alone needs several hours to measure a frequency it does not already know. A fleet that migrates often can therefore spend much of its time not fully converged. ## Step-at-boot as a policy Because a step is a discontinuity, most operators make it a deliberate event: 1. **At boot, step freely.** Before applications start, nothing depends on continuous wall-clock time, so the first correction can be as large as needed. 2. **After boot, slew small errors** and accept RFC 5905's delayed step for large ones, or forbid steps entirely for workloads that cannot tolerate a jump. 3. **Alert on every step after boot.** A step in a running system means migration, suspension, a failing oscillator or a bad source, and each needs a look. 4. **Know the cost of forbidding steps.** At a 500 ppm slew limit, an operating system's limit rather than an RFC rule, an 8 s offset takes 16,000 s, about 4.4 hours, to slew out. A common remedy for migration itself sits outside NTP: the platform corrects the guest clock on resume, or a time synchronisation is triggered as part of the migration, so that the NTP client sees only a small residual. ## The panic threshold and restarts The panic threshold exists so that a client refuses an absurd offset. RFC 8633 section 5.2 describes how that protection disappears under two conditions together: - the operating system restarts the NTP client automatically when it exits, and - the client is configured to ignore the panic threshold on **every** restart. An attacker who can deliver one large forged offset then makes the client exit, and on restart the client accepts that offset. RFC 8633 says operators SHOULD NOT ignore the panic threshold in all cold-start situations without sufficient oversight, and offers mitigations: - monitor the system log for panic exits; - require manual intervention for a step above the panic threshold; - ignore the panic threshold only in a genuine cold start; - require more servers to agree before the clock is adjusted. It also asks implementers to refuse steps to a time earlier than the daemon's build date. Securing the time sources themselves, with authentication and NTS, is a separate subject. ## Common mistakes - Expecting NTP to fix a resumed guest instantly. - Disabling the panic threshold everywhere to make suspended VMs recover. - Assuming only the phase is wrong after migration. - Forbidding steps without budgeting the hours of slewing that follow.

  • A guest resumes two hours behind after a suspension; what does an NTP client following RFC 5905 do?
    7,200 s exceeds the 1000 s panic threshold, so the client SHOULD exit with a log message rather than accept the offset. The safe remedies are a monitored cold-start step allowed past the panic threshold, or having the platform correct the guest clock before the client resumes, not disabling the panic check on every restart.
  • Why does a migrated guest keep drifting after its offset has been corrected?
    The frequency correction it learned describes the source host's oscillator. On the destination host the rate differs, so a fresh frequency error appears, and the discipline must relearn it through its feedback loop. Until it does, offsets build between polls and can cross the step threshold again.

saying these in an interview costs you the question

  • NTP steps the guest clock the moment it sees an offset above 125 ms
  • Disabling the panic threshold on every restart is the safe fix for suspended VMs
  • After migration only the time is wrong; the learned frequency still applies
  • RFC 5905 forbids stepping a clock once the system is running
  • Slewing an eight-second offset completes within a minute