skip to content

How do you choose the poll interval and the timeout for a wait-until-condition helper in a test?

level: middleimportance: must knowfreq 63%

answer

  1. Two numbers, two different questions
  2. Interval is resolution against overhead
  3. Timeout is a claim about the system
  4. Measure a high percentile, add headroom
  5. Report the last observed value

basics

~20 s

Scale the poll interval to the cost of one check: small for a cheap local read, larger than the latency of a network call. Derive the timeout from a measured high percentile of the real wait plus headroom.

solid answer

~50 s

The two numbers answer different questions. The **poll interval** trades resolution against overhead: it should sit comfortably above the cost of a single check, so the helper is not spending its budget on its own queries, and it can back off with a little jitter when the check hits a shared dependency. The **timeout** is an assertion about the system — "if this has not happened by now, something is wrong" — so derive it by measuring the wait's real duration across many runs, taking a high percentile rather than the mean, and adding stated headroom. Three mechanics matter as much as the numbers: evaluate the predicate once before the first pause; decide deliberately whether a thrown check means "not ready" or a hard failure; and make the timeout message report what was awaited and the last value observed, so a timeout is a diagnosis rather than a mystery.

code

pseudocode · 19 lines
pseudocode
function wait_until(describe, check, accept, timeout, interval, max_interval):
    deadline = now() + timeout
    checks   = 0
    last     = "<never observed>"
    while true:
        checks = checks + 1
        try:
            value = check()
            last  = value
            if accept(value):
                return value
        catch NotFoundError:
            last = "absent"          # not ready yet
        # any other error propagates immediately
        if now() >= deadline:
            fail("waited " + timeout + " (" + checks + " checks) for "
                 + describe + "; last seen " + last)
        sleep(min(interval, deadline - now()))
        interval = min(interval * 1.4, max_interval)

go deeper

for a junior

Know that a wait needs both a repeat interval and an overall budget, and that the budget must be long enough to cover a slow but healthy run without being so long that failures stall the suite.

for a middle

Explain the mechanics: check before the first pause, scale the interval to the cost of one check, derive the timeout from measured percentiles, and put the last observed value into the failure message.

for a senior

Show that you tune these from evidence you collected — recorded wait durations, a percentile, stated headroom — and that you treat a wait that routinely spends seconds on the happy path as information about the system rather than as a number to adjust.

for a principal

Frame timeouts as service-level assertions the organisation owns: one shared helper, defaults justified by measurement, overrides that must cite evidence, and a review habit that treats a rising timeout as a performance regression report.

### What the two numbers mean A wait-until-condition helper takes a predicate and repeats it until it holds. Two numbers govern it. The **poll interval** is how long the helper pauses between evaluations of the predicate. The **timeout** is the total budget it will spend before it gives up and fails the test. They answer different questions, and confusing them is the root of most badly-behaved waits. The poll interval controls **resolution and overhead**: it decides how much later than the truth your test learns that the condition became true, and how much load the checking itself puts on the system. The timeout controls **the claim you are making about the system**: "if this has not happened within N, something is wrong". A timeout is, in effect, an assertion about a service level — and it is worth saying that out loud in an interview, because it reframes the number from a magic constant into a statement someone can be held to. ### Choosing the poll interval Start from the cost and the latency of one check. - **Cheap and local check** — reading an in-process value, a single indexed lookup: a small interval, in the low hundreds of milliseconds, is fine. Below roughly a tenth of a second you are mostly burning CPU for resolution nobody perceives. - **Expensive check** — a network round trip, an unindexed query, a call that itself takes 300 ms: the interval should be at least a small multiple of the check's own cost, or the helper spends its whole budget waiting on its own queries and starves the system it is watching. - **Shared or rate-limited dependency** — a poll that hammers a service can change the behaviour it is measuring. This is where backoff (lengthen the interval after each failed check) and a little jitter (randomise it slightly, so parallel workers do not synchronise) earn their keep. Two mechanical details matter more than the number. **Evaluate the predicate once before the first pause**, so a condition that is already true returns immediately instead of paying an interval for nothing. And when the predicate throws — the record does not exist yet, the endpoint is not up — decide deliberately whether that counts as "not ready yet" (swallow and retry) or as a hard failure. Swallowing everything is how a wait silently absorbs a genuine error for its entire budget and then reports a bland timeout. ### Choosing the timeout Derive it from observed behaviour, not from a round number that felt safe. The usable procedure: 1. Measure the wait's real duration across many runs, including a loaded machine and a cold start. 2. Take a high percentile — the 99th, not the mean, because the mean is not what fails you. 3. Add headroom, commonly two to three times that percentile, and record why you chose it. Worked example on a smart-meter reading feed: the wait for a batch's daily total to close measures 1.9 s typical and 6.4 s at the 99th percentile across a few hundred runs. A timeout somewhere near 15-20 s carries clear headroom while still being far below anything a person would call "fine". A 2-second timeout would fail roughly one run in a few dozen for no reason; a 120-second timeout would turn every real breakage into a two-minute stall that tells you nothing. Note the asymmetry, because strong candidates raise it: **a generous timeout costs nothing on a passing run**, since the wait returns the moment the condition holds. It costs on failing runs — and it costs *diagnosis* always, because a timeout so large it can never plausibly trip has stopped being an assertion about the system. ### The failure message is part of the design A timeout that reports only "condition not met in 20s" throws away everything the helper knew. A well-built helper reports: what it was waiting for, in words; the **last observed value** of whatever it was inspecting; how many times it checked; and the elapsed time. "Waited 20.0s (63 checks) for daily total of meter M-4471 to reach window_state=CLOSED; last seen OPEN with 5,120 of 8,412 readings" turns a timeout into a diagnosis. That single habit converts more mystery failures into fixed bugs than any tuning of the two numbers. ### Interval and timeout are not independent of the suite Each wait's budget is spent from a shared pot. A suite of 27 minutes with a few hundred waits can absorb generous per-wait timeouts precisely because they rarely trip — but if several waits routinely spend seconds each on the happy path, that is a signal you are waiting on something slow, not that the numbers are wrong. Also make sure the interval divides the budget sensibly: an interval of 5 s against a timeout of 6 s gives the predicate two chances, so the wait is nearly a coin flip dressed up as a poll. ### What good looks like One shared helper, so the two numbers are set in a single place and tuned with evidence; per-call overrides for genuinely slower operations, with a comment saying which measurement justified the override; the predicate evaluated before the first pause; explicit handling of thrown checks; and a failure message that names the artefact and the last value seen.

  • Does a generous timeout slow down a suite that is passing?
    No — a wait returns the moment its condition holds, so on a green run the timeout is never spent. The cost lands on failing runs, where every timeout is paid in full, and on diagnosis: a timeout large enough that it can never plausibly trip has stopped being an assertion about the system and become a stall. That asymmetry is why the argument for tight timeouts is about signal, not about run time.
  • Your helper swallows every exception from the check and retries. What can go wrong?
    It absorbs genuine errors. A misconfigured endpoint, an authorisation failure or a malformed query all look identical to "not ready yet", so the wait spends its whole budget and then reports a bland timeout with no cause attached. Swallow only the narrow, expected "does not exist yet" case, let everything else propagate immediately, and keep the last error in the timeout message.
  • Why does polling with jitter matter when many workers run in parallel?
    Without jitter, workers that started together stay in phase and hit the dependency in synchronised bursts, which can add latency to the very operation they are waiting for and make the waits look slower than the system is. A small random offset on each interval spreads the checks out. Combined with backoff it also keeps a long wait from generating hundreds of near-useless calls.

saying these in an interview costs you the question

  • Picks a round timeout with no measurement behind it
  • Sleeps before the first check of the condition
  • Polls faster than the check itself completes
  • Swallows every exception as not-ready-yet
  • Reports only "condition not met" on timeout
  • Sets an interval nearly as long as the timeout

context