skip to content

One team's script polls a provisioning call in a tight loop, and now every other tool in the shared account is throttled - why does one caller degrade everyone, and what do you do first?

level: seniorimportance: should knowfreq 40%

answer

  1. the meter is the account, not the script
  2. noisy neighbour inside your own walls
  3. cap the caller before adding retries
  4. starved deployments and console first
  5. separate accounts separate the budget

basics

~20 s

Management API throttling is metered for a scope larger than the script - usually the account, often per region - so one tight loop spends a budget every tool shares. Cap or pause that caller first; adding retries elsewhere only makes the meter worse.

solid answer

~50 s

This is the noisy-neighbour problem with the neighbour inside your own walls. The management API's rate limit is normally counted for the whole account, often split by region and by operation class and only sometimes by identity, so a loop calling many times a second leaves the remainder for deployments, dashboards, drift detection, cost tooling and the console your responders are using. The first action is flow control at the source: pause the job or cap its rate, because that is the only step that returns budget to everyone else. Identify the caller from the platform's record of management API calls, which carries the identity and the operation behind each request. Then fix it structurally - space the loop, give the job a client-side call budget, and move noisy batch work into its own account so its worst case is bounded by something other than good behaviour.

go deeper

for a junior

Take away the core fact: the request-rate limit belongs to the account, not to your script, so a loop you wrote can stop your colleagues' deployments even though it only touched your own resources.

for a middle

Explain the metering scope and the consequence, and be able to say why retrying harder in the affected tools adds load to the very meter that is saturated.

for a senior

Walk the incident in order: confirm the workload is healthy, find the top caller in the record of management API calls, cap it, then fix the call pattern. Mention that the console you are investigating from is throttled too.

for a principal

Frame it as a shared resource with no owner. Decide which callers get permanent caps, which work moves to its own account, and what visibility exists before the next incident rather than during it.

## One budget, many callers The reason a single script degrades everything is arithmetic, not malice. Management API throttling is metered for a **scope larger than the caller**: most commonly the account, usually also per region, frequently split by operation class, and on some platforms additionally per identity. A tight loop calling the same management operation many times a second consumes that shared allowance, and every other client of the API in that scope gets whatever is left. This is the noisy-neighbour problem, except you own both the neighbour and the victims - which is good news, because it means you can fix it without a support case. ## What starves, in roughly this order - **Deployments.** Pipelines describe, create and update constantly, so they fail first and loudest. - **Automation that replaces things.** Anything that reacts to a failure by launching, registering or reconfiguring needs the same API. - **Dashboards and inventory.** Tools that refresh resource state add load while showing you less. - **Cost, compliance and drift tooling.** Sweeps that were already marginal begin failing outright. - **The humans.** The console is a client of the same API, so the responders investigating are themselves throttled. That detail is what makes this incident unusually disorienting. Note what does **not** starve: the workload's own request path. Users are served by the data plane, metered separately. A control-plane throttling event costs you the ability to *change* the estate, not the ability to serve it - unless changing it is exactly what you needed. ## The order of operations 1. **Confirm the shape.** User-facing latency and error rate flat; management calls failing with a throttling response rather than a validation error or a timeout; failures confined to your account. 2. **Find the caller.** The platform's record of management API calls carries the identity and operation behind every request. The top caller by call count over the last few minutes is almost always the answer. 3. **Apply flow control at the source.** Pause the job, or cap its rate. This is the only action that returns budget to everyone else. 4. **Let the other callers recover on their own terms.** They should already be backing off with jitter; if any of them retries immediately, that is the second defect to fix. 5. **Only then change structure.** Nothing structural can be evaluated while the meter is saturated. ## Separating the meters afterwards | Option | What it separates | What it costs | |---|---|---| | A client-side call budget in the job | Nothing, but bounds the damage | Discipline; binds only code you control | | A dedicated identity for the job | The budget, where the platform meters per principal | Little, but buys nothing where it does not | | A separate account for the workload | The budget, on every platform | Another account to set up, reach and inventory | | Removing the calls | The problem itself | Requires the data to be available another way | The honest ordering is: cap first, because it is immediate and it works everywhere, then move the noisiest scheduled work into its own account so its blast radius is bounded structurally rather than by whoever last reviewed the script. ## Why adding retries is the wrong instinct The starved tools are failing, so the reflex is to make them try harder. That reflex adds calls to an already saturated meter. The offender is in a tight loop and is unaffected by anyone else's retries, so it keeps its share while the victims fight over the remainder. Worse, retries triggered by one event tend to synchronise, so a wave of clients all retries at the same instant and re-saturates the meter just as it recovers. The only retry policy that helps under throttling is one that lowers the rate: exponential backoff with jitter, a cap on attempts, and a willingness to fail the run rather than keep hammering. ## The case that should worry you Throttling on a quiet afternoon is an annoyance. Throttling during a real failure is a different problem: recovery actions - launching replacement capacity, changing where traffic goes, promoting a standby - are themselves management API calls, and they queue behind the sweep. That is the argument for capping batch callers permanently rather than only during incidents, and for keeping every scheduled job's call rate visible on a dashboard before the day it matters.

  • Why can retries in the innocent tools make the incident worse?
    Because they add calls to a meter that is already saturated, and the offending loop is unaffected by them. Retries triggered by the same event also synchronise, so clients re-saturate the budget the instant it recovers. Under throttling, the only useful retry policy is one that lowers the aggregate rate: backoff with jitter and a cap on attempts.
  • The offending job is business-critical and cannot simply be paused. What is the middle option?
    Cap it rather than stop it. A client-side call budget of a few calls a second, a longer interval between polls, and a hard limit on concurrency keep it progressing while returning most of the account's allowance. If it must run at speed, move it to its own account so the blast radius is its own tooling rather than everyone's.
  • How do you tell this apart from the account simply needing a higher ceiling?
    Look at whether the call rate is justified by the work. A loop making repeated identical calls, or a sweep re-reading an unchanged estate, is demand you created rather than demand you need. Raise a ceiling only after the call pattern is defensible, otherwise the same loop saturates the higher ceiling a quarter later.

A building with a shared switchboard: one department's auto-dialler keeps every outside line busy, so nobody else can place a call - including the people trying to phone that department to ask them to stop.

saying these in an interview costs you the question

  • Thinks every script gets its own rate budget
  • Increases retries in the starved tools to win the race
  • Declares a platform outage without checking who called
  • Raises a capacity ceiling to fix a request-rate problem
  • Assumes read calls are never counted against the meter
  • Believes user-facing traffic is degraded by control-plane throttling