skip to content

questions

4

In an automated test, what makes a readiness signal worth waiting on rather than a proxy for it?

level: juniorimportance: must knowfreq 74%

answer

  1. Ask what the system itself publishes
  2. Two properties: sufficient and stable
  3. Signal must outlive a single moment
  4. Poll the artefact the assertion reads
  5. Counters belong to one stage only

basics

~20 s

Wait on state that exists only once the work has finished: a persisted record, a status field at its final value, a drained queue. A proxy such as elapsed time or a progress indicator can be true before the work is done.

solid answer

~50 s

A readiness signal is worth waiting on when it has two properties: it becomes true **only after** the work completes, and it **stays** true once true. Persisted records, terminal status transitions and drained queues usually qualify, because the system itself produces them as part of finishing. Proxies fail one of the two tests — an elapsed duration is uncorrelated with completion, a progress indicator says work started rather than finished, and a log line is often written before the transaction behind it commits. The strongest rule of thumb is to make the signal and the assertion talk about the same artefact: if the check reads a stored record, wait for that record, not for an upstream counter that only proves an earlier stage ran. When the system publishes nothing that means done, that is a gap in the system to raise, not a gap to paper over with a duration.

code

pseudocode · 12 lines
pseudocode
submit_readings(batch_id = "B-8412", count = 8412)

# poll the artefact the assertion will read
total = wait_until(
    check    = lambda: daily_total_for(meter = "M-4471", day = "2026-09-02"),
    accept   = lambda t: t != null and t.window_state == "CLOSED",
    timeout  = seconds(20),
    interval = millis(320),
    describe = "daily total for meter M-4471 to reach window_state=CLOSED"
)

assert total.kilowatt_hours == 41.7

go deeper

for a junior

Be ready to name concrete signals — a stored record, a status field that reached its final value, an acknowledged queue position — and to say why a duration is not one of them.

for a middle

Explain the two properties a signal needs: it is true only after completion, and it stays true. Show that you can spot a signal owned by a different stage than the one you are asserting on.

for a senior

Demonstrate that you trace the signal back to the artefact under assertion, and that you treat a system with no completion signal as a defect in the system's observability rather than as a testing inconvenience.

for a principal

Own the position that readiness signals are a product contract, not a test trick: argue for systems publishing terminal state and completion events so that tests, operators and downstream consumers all ask the same question the same way.

### The problem a wait is actually solving An automated test that triggers asynchronous work has no way to know, from the call it just made, whether the work is finished. The trigger returns immediately; the effect lands later, somewhere else. A **wait** is the test's answer to "is it done yet?", and the whole quality of that wait comes down to one choice: **what state you interrogate**. That state is the *readiness signal*. Everything else — the poll interval, the timeout, the helper's name — is plumbing around it. A readiness signal is worth waiting on when it is **true only after** the work you are about to assert on has completed, and **stays true** once it becomes true. Those two properties, sufficiency and stability, are what separate a signal from a proxy. ### Signals that carry the guarantee Three families do this well: - **A persisted record.** The row, document or object the work was supposed to create exists in the store the assertion will read from. If the assertion reads the same store, the signal and the assertion cannot disagree. - **A terminal status transition.** A field the system itself owns has moved to a value that means "finished" — accepted, settled, published, failed. The system is telling you it is done, in its own vocabulary, and it will not move back. - **A drained work queue or a consumed offset.** The backlog the work sits in is empty, or the consumer's position has passed the item you enqueued. This one needs care: empty can also mean "not yet delivered", so it is only sufficient when combined with evidence the item was accepted in the first place. ### Proxies that look like signals and are not - **Elapsed time.** A duration is a guess about the system, restated as a fact. It is not correlated with completion at all — it is correlated with how loaded the machine was on the day the number was chosen. - **A progress or activity indicator.** These typically say *work started* or *work is happening*, not *work finished*. They also flicker: they can be false, then true, then false again, so a poll can miss them entirely. - **A log line.** Emitted at the moment a stage begins, or before the transaction that stage opened has committed. A test that reads a log is asserting on narration rather than on state. - **A counter that belongs to an earlier stage.** This is the subtle one, and it is the most common cause of a wait that passes ninety-nine times and fails on the hundredth. ### A worked example Consider a suite for a smart-meter reading feed. The pipeline has two stages: an ingest stage that accepts a batch of readings and increments an `accepted` counter, and an aggregation stage that turns those readings into a per-meter daily total. A test submits a batch of 8,412 readings, waits until `accepted` reports 8,412, then asserts that the daily total for one meter equals a known figure. The wait looks disciplined. It polls a real number produced by the real system. But `accepted` is the **ingest** stage's signal, and the assertion reads the **aggregation** stage's output. The test is relying on an ordering assumption — that aggregation always finishes before the assertion runs — that nothing in the system guarantees. On an unloaded machine it holds. When the aggregation stage is briefly behind, the test reads a partial total and fails, and the failure looks random because nothing in the test names the stage that was late. The fix is not a longer wait. It is to move the signal onto the artefact under assertion: poll until the daily-total record for that meter exists and its own status says the aggregation window is closed, then assert its value. Now the signal and the assertion are about the same thing, and the ordering assumption disappears because the test no longer makes one. ### When no good signal exists Sometimes the system genuinely publishes nothing that means "done". That is a finding, not a test problem — it usually means the system is also hard to operate, because whoever runs it in production has the same question. The durable fixes are to have the system expose a status or a completion event, or to give the test a supported query that answers the question directly. Reaching for a duration because the system is opaque converts a missing capability into a permanent source of flake spread across every test that touches the feature. ### The cost side A good signal is not free. Polling a store adds queries; polling a status endpoint adds calls. But a poll that returns as soon as the condition holds costs, on a healthy run, roughly the real latency of the work — which is the floor you cannot beat anyway. A duration-based wait costs its full length on *every* run, including the fast ones, and still fails on the slow ones. That is the trade the whole practice turns on: the good signal is usually faster **and** more reliable, which is why interviewers treat "what would you wait on?" as a screening question rather than a specialist one.

  • A drained work queue sounds like a perfect readiness signal. When is it not sufficient?
    An empty queue is ambiguous: it can mean the item was processed, or that it was never enqueued, or that it was rejected before it arrived. On its own it proves nothing about your item. It becomes usable when paired with evidence the item was accepted first — an acknowledged enqueue, a consumer offset that has passed your item's position — or replaced outright by a signal on the item itself.
  • The system exposes no status field and no completion event. What do you do?
    Treat it as a finding and say so. Whoever operates the system in production has the same unanswered question, so the durable fix is to have it publish a status, a completion event, or a supported query that answers "is this item finished?". As a stopgap, wait on the nearest artefact the assertion itself reads rather than on a duration, and record the missing capability so it is fixed once instead of worked around in every test that touches the feature.
  • Why is a signal that can flip back to false dangerous even if it is genuinely produced by the system?
    A poll samples; it does not observe continuously. A signal that is true for a short window can be missed entirely between two samples, so the wait times out on a run where the work actually succeeded. Worse, the failure rate depends on the poll interval, which makes it look like an unrelated timing bug. Prefer a signal that latches — a terminal status, a stored record — over one that pulses.

A parcel tracker that says "out for delivery" tells you a van left a depot; only "delivered, signed for" tells you the parcel is at the door. Waiting on the first is guessing, waiting on the second is knowing.

saying these in an interview costs you the question

  • Treats a duration as equivalent to a condition
  • Waits on a log line written before the commit
  • Polls a counter owned by an earlier stage
  • Assumes an empty queue proves the item finished
  • Calls a progress indicator proof of completion
  • Lengthens the wait instead of changing the signal

context

open as a page

How do you choose the poll interval and the timeout for a wait-until-condition helper in a test?

level: middleimportance: must knowfreq 63%

basics

~20 s

Scale the poll interval to the cost of one check: small for a cheap local read, larger than the latency of a network call. Derive the timeout from a measured high percentile of the real wait plus headroom.

open as a page

A suite's wait timeouts were all raised to 90 seconds to stop intermittent failures. What does that hide, and what should the team do instead?

level: seniorimportance: should knowfreq 54%

basics

~20 s

A blanket raise withdraws the claim that the system finishes in a known time. It hides performance regressions, wrong readiness signals and real races alike. Measure each wait's actual duration, set per-operation budgets from the tail, and fix the signal.

open as a page

When is a retry inside an automated test a legitimate model of the system's contract, and when is it masking a race?

level: principalimportance: should knowfreq 44%

basics

~20 s

A retry is legitimate when it mirrors a contract the system really offers — at-least-once delivery, bounded eventual consistency, a published idempotent operation — and the case asserts that guarantee. Otherwise the retry is discarding evidence of a race.

open as a page