skip to content

Your create call timed out at the client, you retried, and now two preview environments exist — what prevents that?

level: seniorimportance: should knowfreq 46%

answer

  1. the timeout lost the answer, not the request
  2. name the request, not just the resource
  3. same identifier, same response
  4. generated once, reused on every attempt
  5. a fresh identifier per attempt defeats it

basics

~20 s

A caller-supplied request identifier sent with the create. The timeout lost the response, not the request — the platform had already built one. A retry carrying the same identifier is recognised and returns the original result instead of building again.

solid answer

~40 s

A client timeout says nothing about what the server did; the request may have arrived, been executed, and answered into a connection that was already gone. So the retry has to be the **same request**, not a new one. The mechanism is a caller-generated request identifier sent as part of the create: the platform stores it alongside the result for a bounded window, the first call does the work, and a later call presenting the same identifier gets the original response back, including the original resource's identifier. If that identifier arrives with a different body, a well-built API rejects it as a conflict rather than guessing which request you meant. Where no such identifier is offered, fall back to uniqueness — a deterministic resource name, so the second create collides instead of duplicating.

code

http · 21 lines
http
POST /v1/environments HTTP/1.1
Host: platform.internal
Client-Request-Id: build-7741-create-env
Content-Type: application/json

{"name": "preview-4821", "size": "small"}

HTTP/1.1 202 Accepted
{"id": "env-4821", "state": "creating"}

--- the response above never reached the client; it timed out and replayed ---

POST /v1/environments HTTP/1.1
Host: platform.internal
Client-Request-Id: build-7741-create-env
Content-Type: application/json

{"name": "preview-4821", "size": "small"}

HTTP/1.1 200 OK
{"id": "env-4821", "state": "creating", "replayed": true}

go deeper

for a junior

Take away the core fact: a timeout means the answer was lost, not that nothing happened. Retrying without saying it is the same request can create a second resource.

for a middle

Explain the mechanism end to end — a caller-generated identifier sent with every attempt, stored by the platform with its result, and replayed instead of re-executed.

for a senior

Show the failure you have seen: an identifier minted per attempt, or held only in memory, looks like the mechanism and provides none of it. Name the fallbacks when the API offers no identifier.

for a principal

The question you own is what your own platform's create endpoint promises callers, since every team's automation inherits it — and what the organisation does about duplicates created before it existed.

## A timeout is a statement about your client The single most expensive misreading in platform automation is treating a client-side timeout as evidence that nothing happened. It is evidence that **no answer arrived**, which is compatible with four different realities: | What actually happened | What your client saw | State on the platform | |---|---|---| | The request never arrived | a timeout | nothing exists | | It arrived, was rejected, the response was lost | a timeout | nothing exists | | It arrived, was accepted, the response was lost | a timeout | a resource is being built | | It arrived and is still being processed | a timeout | a resource is being built | Two of the four leave a resource behind. A retry that is treated as a brand new request therefore builds a second one — and on an internal platform that stands up a full preview environment per request, the duplicate is not a stray record but a whole environment, with everything inside it running and charging. ## Naming the request, not just the resource The fix is to give the **request** an identity, so the platform can tell a repeat from a new instruction: 1. The caller generates a request identifier **before the first attempt** — one value for this unit of work. 2. Every attempt, including the first, sends that identifier with the create. 3. The platform records the identifier together with the result it produced. 4. A later create presenting a known identifier is not executed again: the stored result is returned, so the caller receives the original resource identifier as if the first response had arrived. 5. The same identifier arriving with a **different body** is a conflict. Serving the stored result would answer a question nobody asked, and applying the new body would silently reinterpret what the identifier claimed. From the caller's side the retry becomes harmless, which is the whole point: the operation is safe to send again because sending it again cannot produce a second thing. ## Rules the identifier has to obey - **Generated by the caller, not the platform.** A value the server hands you cannot help, because the failure mode is precisely not receiving the server's answer. - **Generated once and reused.** A fresh identifier per attempt defeats the entire mechanism — every attempt then looks new. This is the most common way the mechanism is implemented and still fails. - **Durable enough to outlive the process.** If the retry can happen in a later run, the identifier has to be stored with the work item, not held in memory. - **Unique per unit of work.** Reusing one identifier across genuinely different requests is worse than having none: the second request silently returns the first one's result. - **Bounded in lifetime.** Platforms keep the record for a window, not forever. A retry long after that window is a new request again, so very late retries need a different safeguard. ## When the API offers no request identifier Not every management operation accepts one, and providers differ on which do. The fallbacks all work by making the **resource** unique rather than the request: - **Deterministic naming.** Derive the resource name from the work item, so a duplicate create collides with the platform's own uniqueness rule and fails loudly instead of succeeding twice. - **Read before retry.** Look the resource up by that name before retrying. It closes most of the window, but not all of it: the first call can still be mid-flight while you look. - **Reconcile afterwards.** Run a sweep that finds more than one resource for the same work item and removes the extra. This is a safety net, not a design. ## What this does not buy you - **It is not a guarantee of success.** A replayed response can be the original failure, faithfully returned. - **It covers one call, not a workflow.** A sequence of five creates needs its own record of how far it got; making each call individually safe does not make the sequence resumable. - **It is not a retry policy.** How many times to retry and how long to wait between attempts is a separate question with its own answer. This mechanism decides whether a retry is *safe to send at all* — which is the question that has to be settled first. - **It does not clean up what already happened.** If a duplicate was created before any of this was in place, deleting it is still an explicit call someone has to make.

  • The retry presents the same request identifier but a slightly different body. What should the API do?
    Refuse it as a conflict. Returning the stored result would answer a request the caller did not make, and applying the new body would quietly redefine what the identifier stood for. Either way a caller bug becomes invisible. A conflict is the only response that tells the caller its own retry path is broken.
  • Why must the identifier be generated before the first attempt rather than per attempt?
    Because it has to survive the failure. If each attempt mints its own value, every attempt looks like a new request and the platform builds again — the mechanism is present and useless. Generate it with the work item and store it, so any later attempt, even in another process, presents the same value.
  • What if the management API accepts no request identifier at all?
    Make the resource unique instead of the request: derive a deterministic name from the work item so a duplicate create collides with the platform's own uniqueness rule. Read before retrying to close most of the remaining window, and back both with a periodic sweep that finds two resources for one work item.

A ticket number at a collection counter: hand back the same ticket and the clerk looks up your existing order, while asking again without it starts a second one.

saying these in an interview costs you the question

  • Says a client timeout means the request never reached the platform.
  • Generates a new request identifier on each retry attempt.
  • Believes any create is naturally safe to repeat.
  • Thinks waiting longer between retries prevents duplicate resources.
  • Assumes the platform silently deduplicates identical requests by itself.
  • Reuses one identifier across different requests to be extra safe.