skip to content

A service measures how long an operation took by calling a wall-clock/time-of-day API (e.g. the equivalent of `System.currentTimeMillis()`) before and after the operation, then subtracting. Under what circumstances can this produce a negative or wildly wrong duration, and what's the correct fix?

level: middleimportance: must knowfreq 55%

answer

  1. monotonic = elapsed time, non-decreasing
  2. wall/time-of-day = calendar time, can step backward
  3. System.nanoTime() vs System.currentTimeMillis()
  4. CLOCK_MONOTONIC vs CLOCK_REALTIME on POSIX
  5. monotonic values not comparable across machines/reboots

basics

~20 s

If the system's clock gets adjusted backward (say by a time-sync correction) while you're timing something, the 'end' reading can look earlier than the 'start' reading, giving a negative or nonsense duration. Use a monotonic clock instead - one guaranteed to only ever move forward.

solid answer

~50 s

Time-of-day (wall) clocks represent calendar/civil time and can be stepped backward or forward at any moment, by NTP corrections, manual admin changes, leap-second handling, or a VM pausing during live-migration or suspend, because their job is answering 'what time is it,' not 'how much time has elapsed.' Sampling one before and after an operation means an intervening backward step makes the second sample smaller than the first, producing a negative or understated duration. Monotonic clocks, like `System.nanoTime()` in Java or `CLOCK_MONOTONIC` on POSIX systems, are guaranteed non-decreasing and are the correct tool for measuring elapsed time, timeouts, and intervals; they're not tied to any calendar epoch and generally aren't comparable across machines, processes restarting the counter, or reboots. Wall clocks remain correct for logging/timestamping events that need to correlate to human or cross-machine calendar time. The two must not be conflated: use monotonic for 'how long,' wall-clock for 'when.'

go deeper

for a junior

Should recognize that subtracting two wall-clock readings can go wrong if the clock changes in between, and know 'use a different clock for timing' is the fix, without needing the exact API names.

for a middle

Should name the correct monotonic APIs (e.g. System.nanoTime(), CLOCK_MONOTONIC), explain why wall clocks can step backward, and articulate the 'monotonic for elapsed, wall-clock for calendar time' rule.

for a senior

Should identify realistic production failure modes (timeout logic silently defeated, latency-metric spikes correlated with resync events) and know the limits of monotonic guarantees (not comparable across machines/reboots, platform quirks around sleep/migration).

for a principal

Should treat this as a fleet-wide engineering convention to enforce (code review rule, lint check, or wrapper library) rather than a one-off fix, and reason about edge cases like VM migration affecting even monotonic clocks on some platforms.

## Two conceptually different kinds of clock Operating systems typically expose two conceptually different kinds of clock, and conflating them is one of the most common sources of subtle timing bugs in production software. | Clock | What it reports | What it is for | |---|---|---| | **Time-of-day** (or "wall") | Reports the current calendar/civil time, such as milliseconds since the Unix epoch. | Its entire purpose is to answer "what is the date and time right now," correlated with the outside world. | | **Monotonic** | Reports elapsed time since some arbitrary, unspecified reference point (often boot time, but implementation-defined). | Its entire purpose is to answer "how much time has passed," with the single hard guarantee that successive reads never decrease. | ## Why wall clocks are corrected from outside The reason this distinction matters mechanically is that wall clocks are, by design, subject to external correction. - **NTP synchronization** can slew the clock gently or step it abruptly if the local clock has drifted too far from the reference. - **A system administrator** can manually change the date. - **A leap second** can be inserted or smeared into UTC. - **Virtualized or containerized environments**: a VM that is live-migrated, paused, or resumed from suspend can have its wall clock reset or catch up in a jump when it resumes. None of these adjustments are bugs - they're the wall clock correctly doing its job of tracking real calendar time - but they mean a wall-clock read taken before an operation and another taken after can, in principle, have the second value smaller than the first if a backward correction happened in between. Subtracting start from end in that scenario yields a negative duration, or in a smaller-magnitude case, silently understates or overstates the true elapsed time without an obviously wrong sign to catch it. ## What a monotonic clock guarantees, and what it gives up Monotonic clocks solve this by explicitly not tracking calendar time at all. On POSIX systems this is `CLOCK_MONOTONIC` (or `CLOCK_MONOTONIC_RAW` for a version unaffected even by NTP's rate-slewing adjustments); in Java it's `System.nanoTime()`; most language runtimes expose an equivalent. These clocks are guaranteed to be non-decreasing for the lifetime of the process reading them, which is exactly the guarantee an elapsed-time measurement needs. The trade-off: - A monotonic clock's absolute value is meaningless outside the process/machine that produced it - you cannot compare `System.nanoTime()` values across two different JVMs, let alone two different machines. - The clock typically resets on reboot and may or may not continue advancing during system sleep depending on the platform. This makes monotonic clocks unsuitable for anything that needs to correlate with a human-readable date or with events on other machines - that job still belongs to the wall clock. ## The production failure modes The production failure modes from getting this backward are common and often subtle. 1. **Timeout and retry logic** implemented with wall-clock subtraction can compute a negative or near-zero "time remaining," causing either premature timeout firing or, worse, a timeout that never fires because the negative delta is treated as "plenty of time left," silently defeating the timeout's purpose. 2. **Duration-based metrics** (request latency histograms, job-runtime alerting) computed from wall-clock deltas can show impossible negative values or occasional huge spurious spikes that correlate suspiciously with known NTP correction windows or VM live-migration events - a classic on-call debugging trail is "the p99 latency graph shows a few-second spike at exactly the moment ops ran a fleet-wide clock resync." 3. **Rate-limiting or backoff algorithms** that compute "time since last attempt" via wall clock can misbehave the same way. 4. **Scheduled-task frameworks** that compute "time until next run" from wall-clock deltas can fire early, late, or in a tight loop if the wall clock jumps. ## The convention that fixes it The fix is a firm convention, not a one-off patch: - Any code measuring an interval or duration - benchmarking, timeouts, rate limiting, backoff, scheduling relative delays - should use the monotonic clock API. - Any code that needs to record, display, or compare an absolute point in time - log timestamps, "created_at" database columns, cross-service event correlation, expiry dates tied to a real calendar deadline - should use the wall clock. A useful worked example: measuring an HTTP request's latency for a metrics dashboard should be `nanoTime_end - nanoTime_start` (monotonic, immune to clock corrections), while the timestamp attached to the log line describing that request should be `currentTimeMillis()` (wall clock, meaningful to a human reading logs at a specific real time). Many production incident postmortems trace back to exactly this API being chosen for the wrong purpose.

  • Why can't you just use a monotonic clock reading as a substitute for a log timestamp?
    A monotonic clock's zero point is arbitrary and implementation-defined (often boot time), so its value carries no information about the actual calendar date or time and isn't comparable across processes, machines, or even the same machine after a reboot. It answers 'how much time has elapsed since some unspecified moment,' not 'what time is it,' so it can't tell a human or another system what happened when in real, correlatable terms.
  • Does using a monotonic clock for elapsed-time measurement fully protect against clock issues during a VM live-migration?
    Not entirely - some platforms can still cause a monotonic clock to pause or behave oddly across a migration or suspend/resume cycle, since the guarantee of 'never decreasing' is generally solid but the guarantee of 'advances at a real-time rate matching wall-clock elapsed time' is weaker and platform-dependent. It's still far safer than a wall clock for avoiding negative durations, but very precise cross-migration timing may need platform-specific verification.
  • If NTP is slewing the wall clock gently rather than stepping it, is it still unsafe to use for elapsed-time measurement?
    It's safer than a hard step since slewing never runs the clock backward, but slewing does temporarily change the clock's rate, so a wall-clock-based elapsed-time measurement taken during a slew will be slightly inaccurate even though it won't go negative. For anything precision-sensitive, monotonic clocks remain the correct choice regardless of whether the wall clock is being slewed or stepped.

A wall clock is like the calendar on your kitchen wall - it can be corrected if it's wrong, and someone might set it back an hour for daylight saving. A stopwatch is like a monotonic clock - once you press start, it only counts forward, and it's useless for telling you what today's date is, but it's exactly the right tool for timing how long the pasta has been boiling.

saying these in an interview costs you the question

  • Uses wall-clock subtraction for timeout, backoff, or benchmark logic
  • Doesn't know monotonic clock values are meaningless across processes/machines
  • Assumes wall clocks only ever move forward
  • Can't name the platform-specific monotonic API (nanoTime, CLOCK_MONOTONIC, etc.)
  • Suggests fixing negative durations by clamping to zero instead of switching clock source

context