skip to content

Sampling and Wall-Clock Time

Why a profile can show nothing hot while the job is still slow: deterministic profilers count calls, not waiting, and they miss threads. Wall-clock accounting and sampling are the expected next move.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is the difference between wall-clock time and CPU time when timing Python code?

level: juniorimportance: must knowfreq 55%

answer

  1. Two clocks answer two different questions
  2. One counts real seconds, one counts processor seconds
  3. Sleeping and waiting cost only one of them
  4. perf_counter versus process_time
  5. Big gap means waiting, not computing

basics

~20 s

Wall-clock time is elapsed real time, read with time.perf_counter(). CPU time is the time the processor actually spent executing the process, read with time.process_time(). Sleeping, waiting on a socket or blocking on a lock adds wall time but no CPU time.

solid answer

~40 s

`time.perf_counter()` returns a monotonic wall-clock counter in fractional seconds: it counts real elapsed time, including every moment the process spent asleep, blocked on I/O, waiting on a lock or descheduled by the operating system. `time.process_time()` returns CPU time — system plus user time actually consumed by the process — and explicitly excludes time spent sleeping. Subtract two readings of the same clock to get a duration; never mix the two clocks, and never use `time.time()` for durations because it is a wall date-time that can jump backwards when the system clock is adjusted. The gap between the two answers the first diagnostic question about any slow job: a large wall time with a tiny CPU time means the program is waiting, not computing, and no amount of optimising Python code will help.

code

python · 8 lines
python
import time

wall_start = time.perf_counter()
cpu_start = time.process_time()
time.sleep(0.25)               # waiting: costs wall time, no CPU
sum(range(2_000_000))          # computing: costs both
print(f"wall {time.perf_counter() - wall_start:.3f}s")
print(f"cpu  {time.process_time() - cpu_start:.3f}s")

go deeper

for a junior

Be ready to name the two calls and say what each measures: time.perf_counter() for elapsed real time, time.process_time() for processor time used. Also say why time.time() is the wrong tool for durations.

for a middle

Explain the mechanics: perf_counter is monotonic with the platform's best resolution, process_time sums user plus system CPU across the process's threads and excludes sleep. Show that the gap between them classifies a job as waiting or computing.

for a senior

Demonstrate using the ratio as the first diagnostic on a slow production job, before reaching for any profiler, and explain what a CPU time larger than wall time tells you about parallelism and released-GIL C code.

for a principal

Own the argument that wall-clock and CPU accounting should be recorded per phase in long-running jobs by default, so a slowdown is classified from telemetry rather than reproduced locally, and that machine-to-machine wall-clock comparisons are only valid under controlled load.

Timing anything in Python starts with choosing the clock that answers the question you are actually asking, and the `time` module deliberately exposes several because they measure different things. ## Wall-clock time **Wall-clock time** is what a stopwatch on the desk measures: real elapsed time between two moments. In Python you read it with `time.perf_counter()`, which returns a float of fractional seconds from an undefined origin. Only differences are meaningful — the absolute value has no calendar meaning. Two properties make it the right default for measuring durations: * It is **monotonic**: it never goes backwards, so an NTP correction or a daylight-saving change in the middle of your measurement cannot produce a negative duration. * It has the **highest resolution** the platform offers, typically nanoseconds, which matters when the thing you are timing is short. `time.monotonic()` shares the monotonic guarantee; `perf_counter()` is the one documented as the highest-resolution choice for benchmarking. `time.perf_counter_ns()` returns the same clock as an integer number of nanoseconds and avoids float rounding, which is worth using when you accumulate many small intervals. `time.time()` is a different animal: it is the wall *date-time*, seconds since the Unix epoch, and it is not monotonic. It is the correct clock for "when did this happen" and the wrong clock for "how long did this take". ## CPU time **CPU time** is how much processor time the operating system charged to your process — user time in your code plus system time in kernel calls made on your behalf. `time.process_time()` returns it, in fractional seconds, and the documentation is explicit that time elapsed during sleep is *not* included. `time.process_time_ns()` is the integer-nanosecond variant. Because it is a sum over all threads of the process, a program that saturates four cores for one real second reports roughly four seconds of process time. That is not a bug — it is the definition. If you want the CPU consumed by one specific thread, that is `time.thread_time()`. ## Where the two diverge The interesting number is the *gap*: ```python import time wall = time.perf_counter() cpu = time.process_time() time.sleep(0.25) # waiting: wall grows, CPU does not sum(range(2_000_000)) # computing: both grow print(time.perf_counter() - wall, time.process_time() - cpu) ``` Wall time that greatly exceeds CPU time means the process spent its life *waiting*. The usual causes are network or disk I/O, a `time.sleep()` in a retry backoff, blocking on a lock or a queue, waiting on a child process, or simply being starved of CPU on a loaded machine. CPU time close to wall time means the process was genuinely computing, and profiling Python code is then the productive next step. The reverse — CPU time much larger than wall time — means real parallelism: several threads running C code that released the GIL, or a free-threaded build (officially supported since Python 3.14) running Python bytecode on several cores at once. ## Why this is the first thing to measure This distinction is what makes a CPU-oriented call profile misleading on an I/O-bound job. A deterministic profiler records function calls and returns; if your program is blocked inside one call for thirty seconds, the profiler faithfully reports one call, and nothing looks "hot". Knowing that wall time is thirty seconds while CPU time is one second tells you immediately that the answer is not in the function-call report at all, and sends you to wall-clock accounting instead — bracketing each phase of the job with `perf_counter()` readings, or using a sampling profiler that samples on wall-clock. ## Practical rules * Measure durations with `perf_counter()` (or `perf_counter_ns()`), never `time.time()`. * Record both wall and CPU for anything long-running; the ratio is free diagnostic information. * Subtract readings of the *same* clock; the clocks have unrelated origins. * Do not compare a wall-clock number from one machine with one from another under different load — wall time includes everybody else's work on that box. * `time.get_clock_info('perf_counter')` reports that clock's resolution and monotonicity on the current platform if you need to know what you are actually resolving. ## Which clock for which question * *How long did this take?* — `time.perf_counter()`. * *When did this happen?* — `time.time()`. * *How much processor did this cost?* — `time.process_time()`. * *How much processor did this one thread cost?* — `time.thread_time()`. Nothing here answers *where* the time went inside the program; that is what phase timers and profilers are for. But the wall-versus-CPU split decides which of those instruments can possibly help, and it is two lines of code, so it belongs at the top of every investigation into a slow Python program rather than after a wasted afternoon of reading call reports.

  • Why is time.time() a poor clock for measuring how long an operation took?
    `time.time()` reports the system date-time, which is not monotonic: NTP corrections, a manual clock change or a leap-second smear can move it backwards or forwards mid-measurement, producing durations that are wrong or even negative. It also has lower effective resolution than `time.perf_counter()`. Use `time.time()` to record when something happened and `time.perf_counter()` to record how long it took.
  • What does it mean when time.process_time() reports more seconds than the wall-clock elapsed?
    That the process used more than one core in parallel, because `time.process_time()` sums CPU time across all threads of the process. It happens when threads run C code that released the GIL, or on a free-threaded build — officially supported since Python 3.14 — where several threads execute Python bytecode simultaneously. It is a signal that parallelism is real, not a measurement error.
  • How would you decide, in one minute, whether a slow script is I/O-bound or CPU-bound?
    Read both clocks around the whole run and compare. If wall time is much larger than `time.process_time()`, the script is waiting — on network, disk, locks or subprocesses — and optimising Python code is wasted effort. If CPU time is close to wall time, the work is real computation and a call profile will find it. The two-line measurement costs nothing and redirects the whole investigation.

Wall-clock time is the taxi meter running while you sit in traffic; CPU time is the fuel actually burned. A long ride that burned almost no fuel means you were parked, not driving.

saying these in an interview costs you the question

  • Using time.time() to measure elapsed durations
  • Believing time.process_time() includes time spent sleeping
  • Assuming a slow program must have a hot function
  • Mixing readings from two different clocks in one subtraction
  • Thinking process_time can never exceed wall time
  • Treating perf_counter's absolute value as a timestamp

context

open as a page

A route-optimisation job runs 40 minutes, yet its Python call profile accounts for two — where did the wall-clock time go?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A deterministic call profile records calls in one thread and counts time inside them, not time spent waiting. Compare time.perf_counter() with time.process_time(), then bracket each phase with wall-clock timers and sample every thread.

open as a page

Why does sys.setprofile installed in the main thread miss work done in worker threads?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The profiling hook set by sys.setprofile is per-thread interpreter state, so it only fires for the thread that installed it. Use threading.setprofile for threads started later, threading.setprofile_all_threads for existing ones, or a profiler that samples every thread's stack.

open as a page

When does time.thread_time() answer a question that time.process_time() cannot?

level: middleimportance: nice to knowfreq 15%

basics

~20 s

time.process_time() sums CPU time across every thread in the process, so one busy thread is hidden among the others. time.thread_time() reports CPU time for the calling thread only, which is how you charge processor cost to one worker.

open as a page