Your deployment tool is rate-limited while the running service it deploys serves user traffic normally - which request rate is being metered?
answer
- two request streams, two meters
- who is the client of the API
- management calls, not user calls
- counted per account, often per region
- throttled describes, flat user latency
basics
~20 sThe management API meters requests per account, separately from the traffic your workload serves. Deployment tools, dashboards and scripts all spend that management budget; user requests do not consume it unless the service itself calls the management API.
solid answer
~50 sA cloud platform meters two different request streams. The **control plane** - the management API through which resources are created, listed, described, changed and deleted - carries its own rate limit, counted usually per account, often per region, and commonly with a larger allowance for reads than for writes. The **data plane** is the traffic your workload itself serves, bounded by the capacity you provisioned rather than by that limit. Your deployment tool, dashboard refreshes, inventory scripts and infrastructure tooling are all clients of the management API, so they share one budget; a user request to your service does not spend it unless your service also calls that API. That is why this failure looks the way it does: `describe` and `create` calls being refused while user latency and error rate stay flat, which points at a caller inside your own account rather than at a platform outage.
go deeper
Remember that the interface used to create and inspect infrastructure is a separate service from the one serving your users, and it has its own request-rate limit that every tool in the account shares.
Explain the meter's dimensions: usually per account, often per region, generally a larger allowance for reads than writes. Then explain why user traffic is unaffected - it is bounded by provisioned capacity, not by that budget.
Show the diagnosis. Flat user error rate plus failing management calls confined to your account means look for your own noisy caller, not a platform incident, and name what you have actually lost: creating, scaling, replacing.
The interesting question is what depends on that budget during a failure. Recovery automation is a management API client too, so a permanently uncapped batch caller is a hidden dependency in your recovery path.
## Two request streams, two meters A cloud platform exposes two different surfaces, and it counts requests to them separately. The **control plane** is the management API - the single interface through which resources are created, listed, described, changed and deleted. Everything that manages infrastructure is a client of it: the web console, the command-line client, the deployment pipeline, the declarative infrastructure tool, dashboards that refresh resource state, cost and inventory scripts, drift detectors and third-party agents. They all authenticate as some principal in your account and they all draw on the same allowance. The **data plane** is your own workload doing its job - serving user requests, reading and writing records, moving bytes. Its capacity is whatever you provisioned or whatever the managed service scales to. It is not counted against the management API's request-rate limit. So when the deployment tool starts receiving throttling responses while user-facing latency and error rate stay flat, nothing about the workload is saturated. The **management** budget is. ## What the management rate limit is counted against Providers differ in the exact dimensions, but the shape is consistent: - **Per account** - the usual scope, which is why every tool in the account shares one budget. - **Per region** - the same account is often metered separately in each region, so identical calls can be throttled in one and succeed in another. - **Per operation class** - reads such as list and describe commonly get a far larger allowance than writes such as create, delete and modify, and some expensive operations carry their own low ceiling. - **Sometimes per principal** - some platforms also count per identity, which is the only thing that makes a separate credential worth anything here. - **Refilled over a window** - the budget replenishes continuously or per interval, so a burst can fail while the same total spread over a minute succeeds. None of those dimensions belong to the workload. Doubling the machines serving users does not raise the management budget, and it can lower your effective headroom, because more machines usually means more automation describing them. ## Telling the look-alikes apart | What you see | What it is | The tell | |---|---|---| | Management calls fail, users unaffected | Control-plane throttling | The same calls succeed once the caller slows down | | Users see errors, management calls fine | Data-plane saturation | The workload's own capacity signals are at their ceiling | | Both fail, for everyone | A platform incident | Other customers and the provider's status reporting agree | | A call is refused regardless of rate | A capacity ceiling, not a rate limit | Waiting and slowing down change nothing | That last row is worth naming explicitly. A **request-rate limit** is about how fast you are calling; a **capacity ceiling** is about how much you are allowed to have. Backing off fixes the first and does nothing at all for the second, and mistaking one for the other wastes the first twenty minutes of an incident. ## Why the platform splits them this way 1. **The management API is shared infrastructure.** It serves every customer in the region at once, so an unbounded caller threatens everyone's ability to manage anything. 2. **Management operations are asymmetric.** One short create call can commit substantial work behind it, so requests cannot be treated as uniformly cheap the way a data-plane read can. 3. **It protects the thing you care about most.** Because the meters are separate, automation that goes haywire degrades your ability to *change* the estate rather than your ability to serve it - bad, but recoverable. ## Using the signal - Confirm the workload really is healthy from the user's side before declaring an incident. Throttled management calls with a flat user error rate is not an outage. - Look inside your own account first. Most control-plane throttling is a caller you own: a loop, an estate-wide sweep, a dashboard refreshing aggressively, or several pipeline runs landing together. - Identify the caller by operation and identity. The platform's record of management API calls carries both. - Fix the call volume before the retry policy. Retrying harder against a saturated meter turns a short episode into a long one. - Remember precisely what you have lost: creating, scaling, replacing and reconfiguring. If an automatic recovery action needs those calls, it is affected too, which is why throttling that coincides with a real failure is far more dangerous than throttling on a quiet afternoon.
- The same automation is throttled in one region and succeeds in another. What does that suggest about how the limit is counted?That the platform meters the management API per account *and* per region, which most do. The account's allowance in the busy region is exhausted while the same account still has headroom elsewhere. It also means moving a sweep to a quieter region relieves the symptom without reducing the total call volume.
- Does adding machines to the workload raise the management API rate limit?No. The limit belongs to the control plane and is unrelated to the capacity you run on the data plane. In practice more machines make it worse rather than better, because registration, health reporting and describe calls from the surrounding automation scale with the fleet while the management budget does not.
saying these in an interview costs you the question
- Thinks the management rate limit also applies to user traffic
- Reads any management API failure as a provider outage
- Assumes each script gets its own rate budget
- Believes scaling the workload raises the management API limit
- Confuses a request-rate limit with a capacity ceiling
- Assumes read calls are free and never throttled