When an NTP client measures its clock offset, when does it slew and when does it step the clock, and why does that matter?
answer
- continuous versus discontinuous correction
- a threshold measured in milliseconds
- 125 ms, then a 900 s wait
- 1000 s means give up
basics
~20 sBelow RFC 5905's 125 ms step threshold an NTP client slews, running the clock slightly fast or slow so time never jumps. In normal operation larger offsets are stepped only after persisting 900 s; beyond 1000 s it should exit.
solid answer
~50 s**Slewing** corrects the clock gradually by running it slightly fast or slow, so time stays continuous and never goes backwards. **Stepping** sets the clock directly, which is instant but makes wall-clock time jump, possibly backwards. RFC 5905 slews any offset below the step threshold `STEPT` (125 ms). In normal operation a larger offset is first treated as a possible spike: the client ignores it and steps only if outliers persist past the stepout threshold `WATCH` (900 s), which resists congestion bursts. After a step, all associations MUST be reset. An offset above the panic threshold `PANICT` (1000 s) SHOULD make the client exit with a log message rather than trust it. At start-up, before the clock has been set, a large offset is stepped at the first update. Slewing is slow: at a 500 ppm rate, an operating system limit rather than an RFC rule, a 125 ms offset takes 250 s.
code
pseudocode · 15 lines// RFC 5905 clock discipline, simplified (section 11.3, Figure 28)
if |offset| > PANICT (1000 s):
log and exit // SHOULD; operator sets the clock
else if |offset| > STEPT (0.125 s):
if state == SYNC:
state = SPIK // first outlier: ignore it
else if state == SPIK and time_since_last_update < WATCH (900 s):
ignore // still waiting out a possible spike
else:
step clock by offset
reset all associations // MUST
state = SYNC
else:
slew: adjust phase and frequency gradually
state = SYNCgo deeper
Know the two words: slewing changes the clock's rate gradually, stepping sets it at once, and stepping can make time jump backwards.
Explain RFC 5905's three thresholds, 125 ms to step, 900 s stepout and 1000 s panic, and walk through why a single large offset is ignored at first.
Show how steps surface in production, out-of-order logs and reset associations, and how long slewing takes for a given offset under an operating system's slew limit.
Weigh continuity against correctness: when an estate should tolerate a brief step and when it should accept hours of slewing, and how that choice is monitored.
## Two ways to correct a clock When an NTP client has computed its clock **offset** (the time error against its chosen sources), it can remove the error in two ways: - **Slew** - change the clock's rate slightly, running it fast or slow until the offset is absorbed. Time keeps moving forward, smoothly, and no reading is ever repeated. - **Step** - set the clock to the corrected value at once. The error disappears instantly, but every process reading wall-clock time sees a discontinuity, forwards or backwards. RFC 5905 (NTPv4) chooses between them with three constants in its clock discipline (section 11.3, Figure 27): | Constant | Value | Meaning | |---|---|---| | `STEPT` | 125 ms | step threshold: below it, slew | | `WATCH` | 900 s | stepout threshold: how long a large offset must persist before stepping | | `PANICT` | 1000 s | panic threshold: above it, do not correct at all | The appendix skeleton code defines `STEPT` as 0.128 s; the specification text and table say 125 ms, and the text is what to quote. ## The state machine in practice The discipline runs a small state machine. In normal operation it sits in `SYNC`. For each update from the chosen source: 1. If the offset exceeds `PANICT`, the update is a `PANIC`: RFC 5905 says this SHOULD make the program exit with a diagnostic message to the system log, so an operator sets the clock by hand. 2. If the offset is below `STEPT`, it is an `ADJ` update: the loop adjusts phase and frequency gradually, and the clock is slewed. 3. If the offset is above `STEPT` while in `SYNC`, the client does not step. It moves to `SPIK` and ignores the update as a probable spike. 4. In `SPIK`, further outliers are ignored until `WATCH` (900 s) has passed since the last accepted update. An inlier in the meantime returns it to `SYNC` with no step. 5. If the offset is still above `STEPT` after `WATCH`, the clock is **stepped** and the client returns to `SYNC`. At start-up the rule differs: before the clock has been set, an offset above `STEPT` is stepped at the first update, with no spike wait. In normal operation, RFC 5905 says this design resists clock steps under extreme network congestion: a single burst of delayed packets produces a large apparent offset that is not a real clock error, and stepping on it would move a good clock to a bad time. ## What a step costs A step is not just a jump in the displayed time. RFC 5905 states that after a step all peer data are invalid, so **all associations MUST be reset** and the client begins as at initial start; the appendix skeleton also drops the poll interval back to its minimum so it reconverges quickly. Outside NTP, a backward step means wall-clock readings repeat: log lines appear out of order, a file can look older than its source, and anything that subtracts two wall-clock readings can see a negative interval. How applications should measure elapsed time regardless of steps is a distributed-systems and application-design concern, not the protocol's; the protocol's contribution is to step rarely. ## How long slewing takes Slewing trades speed for continuity. The rate at which a clock can be slewed is limited by the operating system or implementation, not by RFC 5905; 500 ppm is a commonly used maximum. At that rate: - 125 ms takes 0.125 / 0.0005 = **250 s**. - 1 s takes 2,000 s, about 33 minutes. - 10 s takes 20,000 s, about 5.6 hours. That is why step thresholds exist at all: below 125 ms slewing is quick and harmless, while a multi-second error slewed out leaves the clock wrong for hours. ## Choosing in operations - **Keep the default step threshold** unless there is a reason. Raising it means more errors are slewed slowly; lowering it means more steps. - **Watch for steps in logs.** A client that steps repeatedly is telling you something is wrong with the oscillator, the virtualisation layer or the sources. - **Treat a panic exit as an alarm**, not as a nuisance to disable; it is the client refusing an offset it has no reason to believe. ## Common mistakes - Saying NTP steps to the server's time on every poll. - Saying NTP never steps a running clock. - Quoting one second as the step threshold. - Believing a panic-sized offset is stepped immediately.
- Why does RFC 5905 wait 900 s before stepping, instead of stepping as soon as an offset exceeds 125 ms?A single large offset is more often a burst of network congestion than a real clock error. The `SPIK` state ignores outliers until the stepout threshold `WATCH` (900 s) passes, so a genuine error is still fixed within about 15 minutes while short bursts never move a good clock. An inlier arriving in the meantime cancels the step entirely.
- What does an NTP client do immediately after it steps the clock?RFC 5905 says all associations MUST be reset and the client begins as at initial start, because every stored sample refers to the old timescale. The appendix skeleton also drops the poll interval to its minimum so the client reconverges quickly after the step.
- Why not raise the step threshold so that NTP never steps a running clock?Then large errors are slewed, which at a commonly used 500 ppm slew limit, an operating-system value rather than an RFC one, takes about 33 minutes per second of offset, leaving the clock wrong for hours. Some deployments accept that for continuity, but it trades a brief discontinuity for a long period of wrong time.
Correcting a pendulum clock that runs fast: you can move the hands, and anyone watching sees the time jump, or you can lengthen the pendulum a little so the clock runs slow until it has caught up, and nobody sees a jump. NTP prefers the pendulum and moves the hands only when the error is large and persistent.
saying these in an interview costs you the question
- NTP sets the clock to the server's time on every poll
- NTP only ever slews; it never steps a running clock
- The NTP step threshold is one full second
- An offset above the panic threshold is stepped immediately
- Slewing a multi-second offset takes only a few seconds