skip to content

Management API Surface

Every resource on a provider is created by a management API call, with the console and command line as clients. Asked because a call that returns immediately does not mean the resource is ready.

on this pageshow

questions

4

Your create call returns an identifier and reports success in under a second, yet the new resource refuses connections — why?

level: juniorimportance: must knowfreq 74%

answer

  1. accepted, not finished
  2. a handle, not a running thing
  3. the control plane took the job
  4. states run until available or failed
  5. success at second zero, ready at minute five

basics

~20 s

The call was accepted, not completed. A management API create is asynchronous: it validates and records the request, hands back an identifier immediately, and only then works the resource through build states before anything can serve traffic.

solid answer

~40 s

Because a create is accepted, not completed. On a management API, provisioning is asynchronous: the platform validates the request, records it, returns an identifier and a state such as `creating`, and only then does the real work — placing capacity, attaching and initialising storage, programming a network path, starting a managed engine, running its own health checks. A success response means the control plane took the job; it says nothing about the data plane being able to serve. The resource then moves through states until it reaches a terminal one, `available` or `failed`, and only at that point is it worth connecting to. The identifier you were handed is a handle for asking about progress, not proof that anything is running yet.

code

http · 14 lines
http
POST /v1/environments HTTP/1.1
Host: platform.internal
Content-Type: application/json

{"name": "preview-4821", "size": "small"}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "id": "env-4821",
  "state": "creating",
  "statusUrl": "/v1/environments/env-4821"
}

go deeper

for a junior

Hold on to one sentence: the create call was accepted, not completed. The identifier it returned is how you ask whether the work has finished.

for a middle

Explain why providers split the operation — provisioning takes minutes — and name the state machine: transient states in the middle, available or failed at the end.

for a senior

Demonstrate the consequences you have lived with: a failed create leaves a record that must be deleted, and a terminal success state is still not a connection your client can open.

for a principal

Your angle is the contract your own platform hands its callers: what a create returns, how progress is observed, and whether teams can build reliable automation against it without guessing.

## What the response actually promised A management API create is **asynchronous by design**. The request you sent asks the control plane to validate something, record it, and begin building. What comes back — a success status, an identifier, and usually a state field reading something like `creating` — is an acknowledgement that the request was accepted. It is not a report that the work finished, because when the response was written the work had barely started. The reason is arithmetic. Provisioning real capacity takes seconds to many minutes: a machine has to be placed somewhere with room for it, storage attached and initialised, a network path programmed, a managed engine installed and started, a replica seeded from a snapshot, and the platform's own probes satisfied. Holding a client connection open for minutes is unworkable — it breaks through proxies, it ties up resources on both ends, and it gives the caller nothing to reconnect to when the connection drops. So the API splits the operation in two: a short call that accepts the request and returns a handle, and a record you can read as often as you like. ## The states a resource passes through 1. **Accepted.** The request is recorded and an identifier exists. Nothing is running yet. 2. **Validating.** Field values, ceilings and permissions are checked. A rejection here can still turn the whole thing into a failure after the call already returned successfully. 3. **Building.** Capacity is allocated and configured. This is where nearly all of the wall-clock time goes. 4. **Checking.** The platform runs its own probes against what it built before it is willing to call the thing usable. 5. **Terminal.** The state settles on the platform's word for success — commonly `available` or `ready` — or on `failed`. Nothing moves after that without another call. The exact words differ between providers. The shape does not: transient states in the middle, exactly two kinds of terminal outcome at the end. ## Four signals and what each one proves | Signal | What it proves | What it does not prove | |---|---|---| | A success status on the create | the request was well formed and accepted | that anything is being built yet | | An identifier in the response | the platform has a record you can query | that a resource exists in usable form | | State reads `available` | the control plane finished its own work | that your client can open a connection | | A connection that succeeds | the serving path is answering you | that it is warmed, replicated or at full capacity | ## Available is not always reachable Even a terminal success state can precede reachability. The name clients will use may still be publishing. The serving path may be accepting its first connections slowly while caches fill. Your own network route to the resource is a separate thing you configured, and whether your caller is permitted to talk to it is a separate mechanism again. None of these are the create call's business; they are the reasons a script should finish with a functional check and not with a status field. ## When it fails half way - The resource usually lands in a **failed terminal state and stays there**. It is a record, not a rollback, and it often has to be deleted explicitly. - Pieces built before the failure may persist. How much the platform cleans up on its own varies between providers and between services on the same provider. - The **reason** for the failure lives on the resource record, not in the create response you already received and discarded. Keeping the identifier is what makes the diagnosis possible. - Something left in a failed state can still be occupying a name, a ceiling or a charge until it is removed. ## What this means for anything you automate 1. Never treat the create response as completion. Treat it as a receipt. 2. Keep the identifier and the status location from the response; they are the only handles you get. 3. Wait on the resource's own reported state, not on the clock. 4. Distinguish a failed create from a slow one — they need opposite reactions, and only the record can tell you which you have. 5. Log the identifier alongside whatever you were doing, so a human who finds the mess an hour later can look up the same record you were watching.

  • The resource reached a terminal success state but your client still cannot connect — what is left?
    Three things outside the create's remit: the name clients use may still be publishing, the serving path may still be warming up, and your own route to the resource is configuration you own. Whether the caller is allowed to talk to it is a separate mechanism again. That is why a wait should end with a real connection attempt.
  • A create failed half way through. What is left behind, and what does it cost you?
    Usually a record in a failed state, sometimes with pieces that were built before the failure. Platforms differ in how much they clean up, so assume nothing rolled back: read the failure reason off the record, delete it explicitly, and check whether what it holds — a name, a ceiling, a charge — has been released.

Ordering a building's fit-out: the confirmation arrives in seconds with a job number, and the job number is exactly what you quote when you ring up to ask whether the work has finished.

saying these in an interview costs you the question

  • Says a success status on create means the resource is usable.
  • Treats the returned identifier as proof something is already running.
  • Assumes a failed provision always leaves nothing behind.
  • Thinks resources created in the console appear faster than through a call.
  • Believes a fixed wait is equivalent to checking the resource's state.
open as a page

A resource created in a provider's web console is identical to one made from the command line — why?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Both are clients of the same management API. The console is a hosted application that turns a form into the same request the command line tool sends; neither has a private path into the platform.

open as a page

Your provisioning script sleeps ninety seconds after a create call instead of polling — what breaks, and what should it wait on?

level: middleimportance: should knowfreq 52%

basics

~20 s

A fixed sleep is wrong in both directions: too short and the script proceeds against a half-built resource, too long and every run pays the worst case. Wait on the resource's own state until it is terminal, then prove the endpoint answers.

open as a page

Your create call timed out at the client, you retried, and now two preview environments exist — what prevents that?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A caller-supplied request identifier sent with the create. The timeout lost the response, not the request — the platform had already built one. A retry carrying the same identifier is recognised and returns the original result instead of building again.

open as a page