A nightly job that describes every resource in the account to check ownership tags is now throttled - how do you cut its call volume?
answer
- fewer calls, not faster retries
- one call per resource is the defect
- page through, never restart
- ask the list call for the field
- skip what the marker says is unchanged
basics
~20 sStop asking about one resource at a time. Read the fields you need from the paged list call, follow the page token instead of restarting, cache anything whose change marker is unchanged, narrow the scope, and pace the run. Backoff only spreads the calls you still make.
solid answer
~50 sThe job is throttled because its call count grows with the estate, not because the platform got stricter. The usual shape is one list call per page followed by one description call per resource, so an account with thousands of resources makes thousands of management API calls to read a single field. The fixes in order of effect: ask the **list** call for the fields you need so the per-resource calls disappear entirely; request the maximum page size and follow the returned page token rather than restarting after a failure; cache each resource's last-seen change marker and skip anything unchanged; narrow the scope to the resource types and regions that can actually carry the tag. Only then shape the rate - a client-side budget of a few calls a second, plus backoff with jitter on a throttling response. The first group removes calls; the second merely spreads them.
code
pseudocode · 31 lines// before: one list, then one description call per resource
resources = listEverything() // restarted from scratch after any failure
for each r in resources:
detail = describeOne(r.id) // one management API call per resource
checkOwnership(detail.owner)
// after: page, read the field inline, cache, and yield when throttled
pageToken = null
attempt = 0
budget = CallBudget(callsPerSecond = 5)
repeat:
budget.await()
page = list(pageSize = maxPageSize,
token = pageToken,
fields = [id, owner, changeMarker])
if page is throttled:
attempt = attempt + 1
sleep(backoffWithJitter(attempt))
continue // retry the SAME token; no page is lost
attempt = 0
for each item in page.items:
if cache.markerFor(item.id) == item.changeMarker:
continue // unchanged since the last run
checkOwnership(item.owner)
cache.put(item.id, item.changeMarker)
pageToken = page.nextToken
until pageToken is nullgo deeper
Recall that every question asked about a resource is a separate call to the management API, and that a job asking about thousands of resources one at a time will be told to slow down.
Explain the call arithmetic and rank the fixes: reading the field from the paged list call removes calls, while sleeping between calls only spreads the same number over more time.
Show the pacing and caching design, including the change marker that makes the cache safe, and justify letting the job run longer so the deployment pipeline keeps its headroom.
The lasting question is whether every team should be running its own sweep at all, or whether one inventory job should publish the estate and the rest should read it.
## Why an estate sweep is the classic throttling story A job that touches every resource in an account is the one workload whose cost grows with the estate rather than with traffic. It is usually written early, when the account holds a few hundred things, and it is written the obvious way: list what exists, then ask the platform about each item in turn. Nothing about it changes as the account grows into the thousands - except the number of management API calls it makes every night, which grows in step. Eventually the sweep crosses the account's **request-rate limit** on the management API and starts receiving throttling responses. It is not asking for more capacity and it is not holding anything. It is simply calling too fast, too often, for too little. ## The call pattern, drawn out 1. One **list** call per page of results. At a page size of 100 across an estate of 8,000 resources, that is about 80 calls. 2. One **describe** call per resource, to read the one field the sweep cares about - 8,000 calls. 3. Any retry of a failed call, which in a naive client is immediate and therefore lands on an already saturated meter. That is roughly 8,080 calls, of which more than ninety-nine percent exist to read a single field. Assume, purely for the arithmetic, that the account's read allowance is in the tens of calls a second: a client firing as fast as it can is throttled inside the first second, and even a perfectly paced version of the same job needs several minutes of the account's entire read budget. The real allowance differs by provider and by operation class; the shape does not. ## Cutting volume beats cutting speed | Change | What it changes | Effect | |---|---|---| | Read the field from the list call | Removes the per-resource call | Largest: turns one call per resource into one per page | | Use the maximum page size | Fewer list calls for the same data | Proportional to the page size | | Follow the page token, never restart | No repeated work after a failure | Removes duplicated pages | | Cache on a change marker | Only changed resources cost a call | Scales with churn, not with the estate | | Narrow the scope | Fewer resources considered at all | Proportional to what you drop | | Pace the calls | Same total, spread over time | No volume change; protects other callers | | Back off with jitter | Ends a retry storm | No volume change; restores progress | The first four reduce how many calls exist. The last three change when they happen. Only the first group survives the next doubling of the estate, which is why adding a sleep is a patch and no longer describing one resource at a time is a fix. ## The parts worth spelling out - **Ask the broad call for what you need.** Most management APIs return more than an identifier in a list response, and many let you request specific fields. Every field you obtain there is a call you never make. - **Page, do not restart.** A paged listing hands back a token for the next page. If a call is throttled, retry that same token; discarding it and starting again multiplies the work at exactly the moment the account has least budget. - **Cache against a marker, not a clock.** Store each resource's last-seen change marker or version and skip anything unchanged on the next run. Cache the description, never the decision - the rule you apply may have changed even when the resource has not. - **Spread the sweep.** No human is waiting on it. A deliberately low, steady rate, off-peak, leaves headroom for the callers that do have someone waiting. - **Treat throttling as flow control.** A throttling response is the platform telling you your rate is too high. Backing off with jitter, so parallel workers do not retry in lockstep, converts a failed run into a slower successful one. ## What to measure afterwards - Calls made per run, plotted against resource count. The ratio should be flat or falling rather than tracking the estate. - The fraction of calls throttled, which should approach zero once the job is paced. - Wall-clock duration, which is allowed to grow. A sweep that takes an hour and disturbs nobody beats one that finishes in five minutes and starves the deployment pipeline. - Cache hit rate, which tells you whether the change marker is genuinely stable or whether something is touching every resource anyway.
- The sweep shares an account with a deployment pipeline. Why pace it even on nights when it is not being throttled?Because the budget is shared and the pipeline has a human waiting on it. A sweep running at full speed can leave just enough budget to pass its own calls while pushing the pipeline into throttling. Pacing a caller with no deadline is how you reserve headroom for the callers that have one.
- Caching descriptions risks acting on stale data. How do you keep the cache honest?Key each entry on a change marker or version the list call returns, treat a missing marker as a miss, and expire entries so a bug cannot preserve a wrong answer forever. Cache the resource's description, never the verdict you derived from it, so a changed rule re-evaluates everything on the next run.
- Is running the sweep from a second account a real fix?It is a containment measure, not a reduction. The calls still exist and still take the same time; they now draw on a different budget, so they no longer starve the first account's tooling. If the sweep must read the first account's resources, it may even make more calls. Reduce volume first, then separate.
saying these in an interview costs you the question
- Retries the same burst immediately and calls it handled
- Adds parallel workers to finish the sweep sooner
- Describes every resource one identifier at a time
- Restarts the whole listing from the first page after a failure
- Assumes a growing estate earns a bigger call budget
- Caches the verdict instead of the resource description