skip to content

A large Kotlin test suite has accumulated dozens of Kotest `eventually` and `continually` calls, and CI time and flakiness are both climbing. How would you decide where polled assertions belong at all, and what policy would you set for their budgets?

level: principalimportance: nice to knowfreq 22%

answer

  1. deterministic > awaited signal > polled
  2. shared config profiles, not scattered durations
  3. narrow scope + narrow expected exceptions
  4. continually always costs its full window
  5. measure total wait time and per-call headroom

basics

~20 s

Treat polling as a fallback for genuine cross-process asynchrony only. Everywhere the code can expose a deterministic signal — an injected clock, an awaited completion, a synchronous test wiring — remove the wait. For what remains, standardise a small set of shared budgets, keep polled scopes narrow, and track total waiting as a metric.

solid answer

~60 s

My framing is a hierarchy, cheapest first: 1. **Make it deterministic.** Inject the clock, expose a completion future, or run the listener synchronously in tests. A wait you can delete is worth more than a wait you tune. 2. **Wait on a signal.** If the boundary can hand you a latch, a callback or an awaitable handle, await that instead of polling. 3. **Poll with Kotest's `eventually`** — only for effects crossing a process or thread boundary you do not control (a broker, a container, a real HTTP server). For what survives, the policy: - **Two or three shared `eventuallyConfig` profiles**, not per-call magic numbers, so budgets are centrally tunable when CI hardware changes. - **Narrow scope**: poll the smallest observable, assert details on the returned value outside the loop. - **Narrow expected exceptions**, so test bugs fail immediately instead of timing out. - **`continually` windows are rationed** — they always cost their full duration, so they belong on a handful of high-value negative assertions. - **Measure**: total time spent inside polled assertions, and how often each one nearly times out. A poll that regularly succeeds at 90% of its budget is a warning, not a pass.

go deeper

for a junior

Say that waits should be as short as possible and that making the code deterministic beats waiting. Naming eventually and continually correctly is enough here.

for a middle

Give the hierarchy — deterministic, awaited signal, polled — and argue for narrow polled blocks and shared configurations instead of ad-hoc durations.

for a senior

Add the diagnostic and cost angles: narrow expected exceptions, rationed continually, and refusing to fix flakiness by doubling deadlines.

for a principal

Own it as suite policy with measurement: total polled time as a tracked number, headroom per call site as an early warning, quarantine rules for retries, and an explicit position on determinism-versus-realism per test tier.

## Why this becomes a real problem `eventually` is the easiest tool in the box to overuse. It reliably turns a red test green, so it spreads: someone adds it to fix a race, the next person copies the pattern, and within a year the suite has a hundred polls with durations ranging from 500ms to 60s chosen by whoever was debugging that day. The symptoms are predictable — CI wall time creeping up, flakiness that moves rather than disappears, and failures that all look like timeouts regardless of cause. The response is not "ban `eventually`". It is a hierarchy of options with polling near the bottom, plus a small amount of policy for the cases that legitimately need it. ## The hierarchy ### 1. Remove the asynchrony from the test Most polling in a typical suite exists because the code under test made time or dispatch implicit. Common fixes: - Inject the clock rather than reading it globally, so "after 30 minutes" is a value you set, not a wait. - Give the asynchronous component a way to report completion — a returned handle, a callback — instead of forcing observers to guess. - In tests, wire the background worker to run inline where that does not invalidate what you are testing. Each of these deletes a wait entirely. Nothing you can do to a poll's configuration is as valuable. ### 2. Await an explicit signal Where the boundary can tell you it is done, await that. Waiting on a signal is exact; polling is a guess about how long exact would have taken. ### 3. Poll, deliberately What is left is real: a broker actually delivers asynchronously, a container actually takes time to become healthy, a browser actually renders later. Here `eventually` is the right tool and should be used with intent rather than reflex. ## Policy for the remainder **Shared budgets, not scattered numbers.** Define a couple of named configurations — say a fast in-process profile with a short deadline and tight interval, and an infrastructure profile with a longer deadline, an initial delay and a backing-off interval. Every call site references one. When CI moves to slower hardware, you change two values; when someone wants a bespoke 90-second budget, that is a reviewable exception rather than an invisible one. **Narrow the polled scope.** Poll for the one condition that is genuinely asynchronous, take the value `eventually` returns, and assert the rest outside the loop. Wide blocks make real bugs cost the full deadline and blur the failure message. **Narrow the retryable exceptions.** With the default "retry everything", a `NullPointerException` in the test itself looks exactly like a slow dependency. Declaring the expected exception types turns test bugs into instant stack traces and keeps the timeout meaning what it says. **Ration `continually`.** It never finishes early, so every use is a fixed tax on every run. Reserve it for negative assertions that genuinely matter — no duplicate processing, no premature firing — and keep the windows short. Ten `continually(5.seconds)` calls are almost a minute of deliberate idling per build. **Never put triggers inside polls.** This is a review rule, not a preference: a state-changing call inside a polled block executes once per attempt, producing tests whose failures are self-inflicted. ## Measurement — the part teams skip Budgets rot silently because a poll that takes longer still passes. Two signals are worth collecting: 1. **Total time spent in polled assertions per run.** This is the direct cost, and it should be a number someone looks at rather than an emergent property of the suite. 2. **Headroom per call site.** A poll that habitually succeeds at 90% of its deadline is one bad CI day from red. That is the moment to fix the underlying determinism, not to double the deadline. Doubling deadlines is the default reflex and it is almost always the wrong one: it converts an intermittent, informative failure into a slow, permanent cost and defers the diagnosis indefinitely. ## Where reasonable people disagree - **Strict determinism vs realism.** Making everything synchronous in tests can hide genuine ordering bugs that only appear with real dispatch. A common compromise: deterministic unit and module tests, a small number of realistic end-to-end tests where polling is expected and budgeted. - **Retry-on-failure at the runner level.** Some teams accept automatic retries for a quarantined subset to keep the pipeline usable. Defensible only if quarantine is visible and time-boxed; otherwise it silently normalises flakiness. - **One global budget vs per-boundary budgets.** A single number is simpler; per-boundary numbers (broker vs container vs HTTP) are more honest about latency profiles. Either beats ad-hoc. ## The stance to express in an interview Polled assertions are a debt instrument. They are the correct tool at a small number of genuinely asynchronous boundaries, and a smell everywhere else. The engineering work is deleting the ones that should not exist and giving the survivors shared, measured budgets — not tuning durations one failure at a time.

  • A poll that used to pass in 200ms now regularly passes at 4.7 seconds of its 5-second budget. What do you do?
    Treat it as a failing test that has not gone red yet. Something upstream got slower — a new dependency in the path, contention, an N+1 query — and doubling the deadline only hides it. Investigate the latency change first; adjust the budget only after you understand why the number moved and have decided the new latency is acceptable.
  • How do you push back on automatic test retries in CI as a flakiness policy?
    Retries convert information into cost: the race still exists, it now surfaces once in ten runs, and every failure round-trip multiplies build time. If retries are used at all they should apply to an explicitly quarantined, visible list with an owner and a deadline, so the debt is tracked rather than absorbed into the pipeline's normal behaviour.

Polled assertions are like buffer stock in a supply chain: a little absorbs real variability, but growing it is how you stop noticing that the upstream process is unreliable.

saying these in an interview costs you the question

  • Doubling every deadline as the standard response to flakiness.
  • Treating `eventually` as a general-purpose fix rather than a boundary-specific tool.
  • Ignoring that `continually` always spends its full window, so its cost is unconditional.
  • Wrapping whole scenarios — including triggering actions — in a polled block.
  • Assuming a green run means the budgets are healthy, with no visibility into headroom.

context