skip to content

Platform Operations

Running an estate after it exists: provisioning through an API, quotas, management-plane rate limits, forced upgrades and an audit record. Asked because these block a launch, not a design.

on this pageshow

questions

22

Your create call returns an identifier and reports success in under a second, yet the new resource refuses connections — why?

level: juniorimportance: must knowfreq 74%

answer

  1. accepted, not finished
  2. a handle, not a running thing
  3. the control plane took the job
  4. states run until available or failed
  5. success at second zero, ready at minute five

basics

~20 s

The call was accepted, not completed. A management API create is asynchronous: it validates and records the request, hands back an identifier immediately, and only then works the resource through build states before anything can serve traffic.

solid answer

~40 s

Because a create is accepted, not completed. On a management API, provisioning is asynchronous: the platform validates the request, records it, returns an identifier and a state such as `creating`, and only then does the real work — placing capacity, attaching and initialising storage, programming a network path, starting a managed engine, running its own health checks. A success response means the control plane took the job; it says nothing about the data plane being able to serve. The resource then moves through states until it reaches a terminal one, `available` or `failed`, and only at that point is it worth connecting to. The identifier you were handed is a handle for asking about progress, not proof that anything is running yet.

code

http · 14 lines
http
POST /v1/environments HTTP/1.1
Host: platform.internal
Content-Type: application/json

{"name": "preview-4821", "size": "small"}

HTTP/1.1 202 Accepted
Content-Type: application/json

{
  "id": "env-4821",
  "state": "creating",
  "statusUrl": "/v1/environments/env-4821"
}

go deeper

for a junior

Hold on to one sentence: the create call was accepted, not completed. The identifier it returned is how you ask whether the work has finished.

for a middle

Explain why providers split the operation — provisioning takes minutes — and name the state machine: transient states in the middle, available or failed at the end.

for a senior

Demonstrate the consequences you have lived with: a failed create leaves a record that must be deleted, and a terminal success state is still not a connection your client can open.

for a principal

Your angle is the contract your own platform hands its callers: what a create returns, how progress is observed, and whether teams can build reliable automation against it without guessing.

## What the response actually promised A management API create is **asynchronous by design**. The request you sent asks the control plane to validate something, record it, and begin building. What comes back — a success status, an identifier, and usually a state field reading something like `creating` — is an acknowledgement that the request was accepted. It is not a report that the work finished, because when the response was written the work had barely started. The reason is arithmetic. Provisioning real capacity takes seconds to many minutes: a machine has to be placed somewhere with room for it, storage attached and initialised, a network path programmed, a managed engine installed and started, a replica seeded from a snapshot, and the platform's own probes satisfied. Holding a client connection open for minutes is unworkable — it breaks through proxies, it ties up resources on both ends, and it gives the caller nothing to reconnect to when the connection drops. So the API splits the operation in two: a short call that accepts the request and returns a handle, and a record you can read as often as you like. ## The states a resource passes through 1. **Accepted.** The request is recorded and an identifier exists. Nothing is running yet. 2. **Validating.** Field values, ceilings and permissions are checked. A rejection here can still turn the whole thing into a failure after the call already returned successfully. 3. **Building.** Capacity is allocated and configured. This is where nearly all of the wall-clock time goes. 4. **Checking.** The platform runs its own probes against what it built before it is willing to call the thing usable. 5. **Terminal.** The state settles on the platform's word for success — commonly `available` or `ready` — or on `failed`. Nothing moves after that without another call. The exact words differ between providers. The shape does not: transient states in the middle, exactly two kinds of terminal outcome at the end. ## Four signals and what each one proves | Signal | What it proves | What it does not prove | |---|---|---| | A success status on the create | the request was well formed and accepted | that anything is being built yet | | An identifier in the response | the platform has a record you can query | that a resource exists in usable form | | State reads `available` | the control plane finished its own work | that your client can open a connection | | A connection that succeeds | the serving path is answering you | that it is warmed, replicated or at full capacity | ## Available is not always reachable Even a terminal success state can precede reachability. The name clients will use may still be publishing. The serving path may be accepting its first connections slowly while caches fill. Your own network route to the resource is a separate thing you configured, and whether your caller is permitted to talk to it is a separate mechanism again. None of these are the create call's business; they are the reasons a script should finish with a functional check and not with a status field. ## When it fails half way - The resource usually lands in a **failed terminal state and stays there**. It is a record, not a rollback, and it often has to be deleted explicitly. - Pieces built before the failure may persist. How much the platform cleans up on its own varies between providers and between services on the same provider. - The **reason** for the failure lives on the resource record, not in the create response you already received and discarded. Keeping the identifier is what makes the diagnosis possible. - Something left in a failed state can still be occupying a name, a ceiling or a charge until it is removed. ## What this means for anything you automate 1. Never treat the create response as completion. Treat it as a receipt. 2. Keep the identifier and the status location from the response; they are the only handles you get. 3. Wait on the resource's own reported state, not on the clock. 4. Distinguish a failed create from a slow one — they need opposite reactions, and only the record can tell you which you have. 5. Log the identifier alongside whatever you were doing, so a human who finds the mess an hour later can look up the same record you were watching.

  • The resource reached a terminal success state but your client still cannot connect — what is left?
    Three things outside the create's remit: the name clients use may still be publishing, the serving path may still be warming up, and your own route to the resource is configuration you own. Whether the caller is allowed to talk to it is a separate mechanism again. That is why a wait should end with a real connection attempt.
  • A create failed half way through. What is left behind, and what does it cost you?
    Usually a record in a failed state, sometimes with pieces that were built before the failure. Platforms differ in how much they clean up, so assume nothing rolled back: read the failure reason off the record, delete it explicitly, and check whether what it holds — a name, a ceiling, a charge — has been released.

Ordering a building's fit-out: the confirmation arrives in seconds with a job number, and the job number is exactly what you quote when you ring up to ask whether the work has finished.

saying these in an interview costs you the question

  • Says a success status on create means the resource is usable.
  • Treats the returned identifier as proof something is already running.
  • Assumes a failed provision always leaves nothing behind.
  • Thinks resources created in the console appear faster than through a call.
  • Believes a fixed wait is equivalent to checking the resource's state.
open as a page

A resource created in a provider's web console is identical to one made from the command line — why?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Both are clients of the same management API. The console is a hosted application that turns a form into the same request the command line tool sends; neither has a private path into the platform.

open as a page

Your managed database has a weekly maintenance window - what may the provider do inside it, and what does your application see?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A maintenance window is a recurring slot you nominate in which the provider may patch and restart your managed instance. The application normally sees dropped connections and a short interruption, or a failover to the standby replica, not a seamless change.

open as a page

What separates a soft quota you can ask to have raised from a hard limit, and what does hitting each cost?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A soft quota is a provider ceiling an increase request can raise, so hitting it costs lead time. A hard limit is fixed by the platform's design, and the only way past it is an architecture change.

open as a page

Your deployment tool is rate-limited while the running service it deploys serves user traffic normally - which request rate is being metered?

level: middleimportance: must knowfreq 62%

basics

~20 s

The management API meters requests per account, separately from the traffic your workload serves. Deployment tools, dashboards and scripts all spend that management budget; user requests do not consume it unless the service itself calls the management API.

open as a page

Why does a managed database service apply minor patches for you but wait for you to start a major version upgrade?

level: middleimportance: must knowfreq 56%

basics

~20 s

A minor patch stays inside one major version line and is meant to preserve behaviour, so the provider can apply it fleet-wide. A major upgrade changes behaviour the provider cannot test against your workload, so you schedule it, test it and own its rollback.

open as a page

Which provider-side record shows who called which management API to change a service overnight, and what does one entry contain?

level: middleimportance: must knowfreq 62%

basics

~20 s

The platform's audit trail of management API calls, a provider-side stream your application never writes to. Each entry names the calling principal, the time, the source address, the operation, the target resource and whether the call was performed, denied or failed.

open as a page

A failover into the second region is refused at a quota, although the primary region ran the same fleet for months — why?

level: middleimportance: must knowfreq 58%

basics

~20 s

Quotas are counted per account and separately per region, so the standby region still holds the defaults the account started with. Years of small raises in the primary never followed the workload across, and nobody had exercised the second region at full size.

open as a page

Using the management-API audit record, how do you reconstruct the timeline of an unexplained overnight change to a shared service?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Bound the window between the last known-good observation and detection, filter to calls targeting that resource and its dependencies, keep denied and read-only calls too, join them by correlation id, and allow for delivery lag.

open as a page

A nightly job that describes every resource in the account to check ownership tags is now throttled - how do you cut its call volume?

level: middleimportance: should knowfreq 46%

basics

~20 s

Stop asking about one resource at a time. Read the fields you need from the paged list call, follow the page token instead of restarting, cache anything whose change marker is unchanged, narrow the scope, and pace the run. Backoff only spreads the calls you still make.

open as a page

Your provisioning script sleeps ninety seconds after a create call instead of polling — what breaks, and what should it wait on?

level: middleimportance: should knowfreq 52%

basics

~20 s

A fixed sleep is wrong in both directions: too short and the script proceeds against a half-built resource, too long and every run pays the worst case. Wait on the resource's own state until it is terminal, then prove the endpoint answers.

open as a page

Why does a management-API audit trail record the creation of an object store but not each read of an object inside it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Creating the store is a control-plane call and is recorded by default; reading an object is data-plane traffic, which is orders of magnitude more frequent and is usually recorded only if you switch that recording on in advance, per resource.

open as a page

One team's script polls a provisioning call in a tight loop, and now every other tool in the shared account is throttled - why does one caller degrade everyone, and what do you do first?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Management API throttling is metered for a scope larger than the script - usually the account, often per region - so one tight loop spends a budget every tool shares. Cap or pause that caller first; adding retries elsewhere only makes the meter worse.

open as a page

Your create call timed out at the client, you retried, and now two preview environments exist — what prevents that?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A caller-supplied request identifier sent with the create. The timeout lost the response, not the request — the platform had already built one. A retry carrying the same identifier is recognised and returns the original result instead of building again.

open as a page

A deprecation notice gives your managed engine's major version an end-of-support date nine months out - what work does that create?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It converts a date into an engineering project: inventory every instance on that version, find what the new version breaks, rehearse the upgrade on restored data, and schedule a cutover with a way back. Let the date pass and the provider upgrades you on its own schedule.

open as a page

Maintenance promoted the standby and the database was back within a minute, yet a nightly job kept failing for an hour - why?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The platform recovered; the client did not. Pooled connections opened before the promotion stayed in the pool, so every borrow handed the job a dead session, and a cached address plus no reconnect logic kept it failing until the process was restarted.

open as a page

In a quarterly access review, how does the management-API audit record tell you which granted permissions a principal has actually used?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Aggregate the operations each principal performed over a review window and subtract them from the actions its grants allow; the difference is the removal candidate list. The window must exceed the slowest legitimate cycle.

open as a page

How do you keep a soft quota from being discovered mid-incident, given an increase request takes days to land?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Track remaining headroom rather than usage: used against granted, per account and per region, converted into days at the observed growth rate. Alert while more days remain than a raise takes, and file the request as ordinary backlog work.

open as a page

Your provider is retiring the instance family your managed instances run on - how does that differ from an engine version reaching end of support?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

The lifecycle is the same - a notice, a date, a forced action - but the change is underneath rather than inside. A family retirement replaces the machine, so the work is a move plus performance and cost re-checking, with no query compatibility to test.

open as a page

Before you commit to a design that leans on a published per-account ceiling, what do you need to establish about it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Four properties, not just the number: whether it can be raised at all, what scope it is counted in, what it counts — things that exist, things running at once, or actions inside a time window — and how long a raise takes.

open as a page

Many teams share one account's management API rate budget; as the platform owner, how do you stop any single automation from starving the others?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Split the meter and shape the demand. Move workloads into their own accounts so budgets stop overlapping, replace every team's estate sweep with one shared inventory, publish call-rate expectations, and alert on call rate per caller before the platform starts refusing.

open as a page

What must be true of a management-API audit record before you can tell an auditor the account's own administrators could not have altered it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The copy that counts must live where the account's administrators cannot delete or overwrite it, with retention outlasting the time it takes anyone to ask, and stopping the recording must itself alert, so silence is visible.

open as a page