Your provisioning script sleeps ninety seconds after a create call instead of polling — what breaks, and what should it wait on?
answer
- waiting is a loop, not a constant
- terminal state, not elapsed time
- available and failed both end it
- every wait needs a deadline
- a state field is not a connection
basics
~20 sA fixed sleep is wrong in both directions: too short and the script proceeds against a half-built resource, too long and every run pays the worst case. Wait on the resource's own state until it is terminal, then prove the endpoint answers.
solid answer
~50 sA constant cannot track a variable. Provisioning time moves with size, region, load and service, so ninety seconds is either an outage waiting to happen or dead time on every run — and it hides failures completely, because a resource that went to `failed` in ten seconds still gets the full ninety before anything notices. Instead, read the record the create handed you and loop until its state is **terminal**: success means proceed, `failed` means stop now with the reason the record carries. Give the loop an overall deadline so a stuck provision cannot hang a pipeline forever, and widen the interval as you wait rather than reading as fast as possible. Then finish with a real check — the control plane calling something available means its own work is done, not that your client can open a connection.
code
pseudocode · 24 linesdeadline = now() + maxWait
interval = initialInterval
loop:
read = get(statusUrl)
if read failed:
# a bad read is not a bad resource
if now() >= deadline: abort("no status before deadline")
sleep(interval); interval = min(interval * 2, maxInterval)
continue
if read.state == "available":
break
if read.state == "failed":
abort("provisioning failed: " + read.failureReason)
if now() >= deadline:
abort("timed out, last state " + read.state + ", id " + read.id)
sleep(interval)
interval = min(interval * 2, maxInterval)
# control plane is done; now prove the serving path answers
until connect(read.endpoint) succeeds or now() >= deadline:
sleep(interval)go deeper
Learn the habit before the theory: after a create, read the resource's state in a loop until it reaches a terminal one. A sleep is a guess, and it cannot notice a failure.
Explain the loop properly — transient states keep waiting, the success state and failed both stop it, a deadline bounds the whole thing, and the interval widens as you go.
Show the second finish line. Control-plane ready is not data-plane ready, so end the wait with a real connection attempt, and make a timeout report the identifier and last state.
Your concern is what the platform your team publishes promises callers: an observable status, a stated maximum provisioning time, and a failure reason good enough that automation need not guess.
## Why a constant is the wrong instrument A fixed sleep encodes a guess about a number nobody controls. How long a resource takes to provision varies with its size, the service, the region, how busy the platform is, whether a replica has to be seeded and whether the platform is doing anything unusual that day. A single constant has to be wrong in one of two ways, and usually manages both in the same week: - **Too short** — the script proceeds against something half-built and fails in a confusing place, often several steps later, where the error says nothing about provisioning. - **Too long** — every run pays the worst case. Multiply that by a pipeline that stands up an environment per change and the waiting becomes the pipeline. - **Blind to failure** — this is the worst part. A create that failed after ten seconds still consumes the full ninety, and then the script carries on as though it had succeeded, because a sleep has no opinion about outcomes. ## What to wait on instead The create call handed you two things: an identifier, and usually a location for reading progress. Reading that record is the wait. Each read returns the platform's current opinion, and your loop's only job is to classify it: | Observed state | What it means | What the loop does | |---|---|---| | a transient state such as building or pending | work is under way | wait and read again | | the platform's success state | the control plane finished its work | stop waiting, then verify functionally | | failed | the platform gave up, and the record says why | stop immediately and surface the reason | | missing or unrecognised | your read did not land, or the vocabulary changed | read again; do not treat it as either outcome | The distinction that matters is **the state of the resource** against **the outcome of your read call**. A read that errors is a fact about your call. A `failed` state is the platform's verdict about the resource. Conflating them makes a script give up on a healthy provision, or wait patiently on one that died. ## Two kinds of ready There are two separate finish lines, and scripts that only cross the first are the ones that fail intermittently. 1. **Control-plane ready.** The management API says the resource reached its success state. All that asserts is that the platform completed the work it had queued. 2. **Data-plane ready.** Something your client sends actually gets answered — a connection opens, a health path responds, a trivial query returns. Between the two sit the published name, the serving path's own warm-up, and your route to the resource. A wait that ends at the first line and hands over to the next step is the classic source of "it works when I run it by hand". ## The shape of a correct wait 1. Take the identifier and status location from the create response. 2. Compute a **deadline** before the loop starts, from a maximum wait you chose deliberately. 3. Read the record. Classify the state against the table above. 4. On the success state, break out and run one functional check. 5. On `failed`, abort now and report the reason the record carries — never wait out the remaining time. 6. Otherwise wait, widening the interval as the wait goes on, and read again. 7. If the deadline passes, abort with the identifier, the last state observed and the elapsed time, so a human can pick up the same record. Widening the interval matters for a second reason: a tight loop against a management API is a load someone eventually notices, and that is a problem with its own consequences you would rather not cause. ## The ways a wait loop still goes wrong - **No deadline**, so a stuck provision holds a pipeline until a human kills it. - **Treating unknown as ready**, which quietly converts a vocabulary change into a false success. - **Treating a read error as a provisioning failure**, which destroys perfectly good resources on a blip. - **Stopping at the state field**, so the next step is the first thing that ever tried to connect. - **Reporting nothing on timeout**, leaving behind a red pipeline and no identifier to look up.
- How do you tell a transient read error during the wait from a real provisioning failure?By whose statement it is. A `failed` state is the platform's verdict on the resource and is terminal. A read that errors or times out is a fact about your call and says nothing about the resource — read again. Only the deadline, not a single bad read, should end the wait without a verdict.
- What belongs in the error when the deadline expires?The resource identifier, the last state you observed, how long you waited, and where the record can be read. That turns a timeout into something a human can continue investigating, instead of a bare message that invites someone to run the whole thing again blind.
- Why widen the polling interval rather than reading at a fixed fast rate?Because the information arrives slowly. Early reads are cheap and occasionally useful; after a minute the state is unlikely to change between one read and the next. Widening keeps latency low for fast provisions while keeping the total number of calls small for slow ones.
saying these in an interview costs you the question
- Sleeps a fixed interval and calls it a readiness check.
- Loops forever with no deadline and no failure path.
- Treats every non-success state as still in progress, including failed.
- Assumes a terminal success state means the endpoint accepts connections.
- Reads the status as fast as possible because the answer is wanted quickly.
- Aborts the whole provision because one status read timed out.