skip to content

How do you keep a soft quota from being discovered mid-incident, given an increase request takes days to land?

level: seniorimportance: should knowfreq 46%

answer

  1. the remedy is slow, so lead it
  2. usage alone is not the signal
  3. used against granted, per scope
  4. convert the remainder into days
  5. alert earlier than a raise takes

basics

~20 s

Track remaining headroom rather than usage: used against granted, per account and per region, converted into days at the observed growth rate. Alert while more days remain than a raise takes, and file the request as ordinary backlog work.

solid answer

~50 s

The lever is slow, so it has to be pulled early, which means the signal has to lead the failure by more than the provider's turnaround. Chart `used / granted` for each ceiling your design leans on, in each account and region — including the standby and the burst paths you only use in an emergency — and convert the remainder into **days of headroom** at the growth you are actually seeing. Alert when that falls below a comfortable multiple of the raise lead time, not when usage crosses a round number. Then treat the raise as a normal backlog item with an owner and a date, and re-ask periodically because the target moves with growth. Two caveats: a granted raise is permission, not a reservation, and a fixed ceiling cannot be rescued by an alert at all, so classify each ceiling before you instrument it.

code

pseudocode · 16 lines
pseudocode
# one ceiling, in one account and one region
headroom = granted - used
if headroom <= 0:
    return "exhausted: creation is already being refused in this scope"

growthPerDay = max(observedGrowthPerDay, 1)
daysOfHeadroom = headroom / growthPerDay

if raisable is false:
    if daysOfHeadroom < redesignNoticeDays:
        return "fixed ceiling approaching: schedule the restructure now"
    return "watch"

if daysOfHeadroom < raiseLeadTimeDays * 2:
    return "file the increase request now: headroom is shorter than the wait"
return "watch"

go deeper

for a junior

Know that a raise is not instant, so a ceiling must be watched before it is reached. The thing to watch is how much room is left, not how much is in use.

for a middle

Explain the signal: used against granted, computed in the scope the ceiling is counted in, turned into days of headroom, with the alert threshold derived from how long a raise takes.

for a senior

Show that you cover the paths that only run under stress — standby, scale-up, recovery — where the number you need is what the path consumes at full size, and name who owns filing the raises.

for a principal

Decide how much headroom the organisation buys as standing policy, since headroom on a raisable ceiling is free while the same discipline on a fixed one is a scheduled redesign. Set what a recovery plan must evidence before acceptance.

## The lever you need is slow, so the signal must be early A soft quota has exactly one remedy — an increase request — and that remedy runs on the provider's schedule, usually measured in days. Every design decision about monitoring quotas follows from that single fact: **your warning must arrive further ahead of the wall than the remedy takes to land.** A signal that fires when the creation call is refused is not a warning, it is a post-mortem. ## Usage is not the signal; headroom is Most estates chart how many resources they are running. That chart never mentions the ceiling, so it cannot tell you how close you are to it. Three levels of signal, in increasing usefulness: | Signal | What it catches | What it misses | |---|---|---| | Absolute usage over time | growth trends, sudden jumps | says nothing about any ceiling | | Used against granted | how close you are, as a ratio | says nothing about *when* you arrive | | Remaining headroom in days | the date you must act by | assumes growth stays roughly smooth | The third is what you alert on. Compute it per scope — one account, one region, one ceiling — because that is how the ceiling itself is counted, and a ratio averaged across regions hides exactly the region that is about to block you. The threshold follows from the lead time, not from taste. If a raise takes several days end to end, an alert at a couple of days of headroom is already too late to act calmly; a comfortable multiple gives room for the request to be reviewed, partially granted, and re-asked. ## Cover the paths you do not exercise daily The ceilings that block you are rarely the ones on your main dashboard, because those are the ones you have already raised. The ones that bite are the ceilings on the paths that only run under stress: - the **standby scope**, which has to create a full fleet exactly once, during an incident; - the **scale-up path**, where an autoscaling group is allowed to triple but the account ceiling is not; - the **recovery path**, which restores volumes, addresses and network constructs all at once, each against its own ceiling; - **seasonal or launch demand**, which is known in advance and therefore has no excuse for arriving as a surprise. For each of these, the useful number is not what you use today but what the path would consume **at full size**. That is a calculation, not a measurement, and it can be redone on paper whenever the fleet grows. ## Make the raise ordinary work Quota raises fail as a practice when they are nobody's job. What makes them reliable: 1. An **inventory** of the ceilings the design leans on, each recorded with its scope, its granted value and whether it is raisable at all. 2. An **owner** for the inventory, and raises filed as backlog items with a date rather than as heroics. 3. A **record** of what was asked and what was granted, since a partial grant is a checkpoint you must return to. 4. A **periodic re-check**, because a raise granted against a smaller fleet quietly stops being enough as the fleet grows. ## What this practice does not buy you Two honest limits, both worth saying aloud in an interview: - **A raise is permission, not a reservation.** Having a higher ceiling means the platform will no longer refuse you on that ground; it does not set anything aside on your behalf, and a refusal that speaks about availability rather than a ceiling is a different problem with a different remedy. - **An alert cannot rescue a fixed ceiling.** If the number cannot be raised, watching it approach only tells you when the redesign becomes urgent. Classify each ceiling first: raisable ones get a headroom alert, fixed ones get an engineering notice far enough ahead that the restructure can be scheduled. ## The interview version of this answer Say the mechanism, then the number, then the ownership. The mechanism is that the remedy has a lead time and so the signal must lead it. The number is days of headroom per scope, alerting at a multiple of that lead time. The ownership is an inventory somebody maintains, covering the emergency paths as well as the daily ones. A candidate who answers only *we monitor our quotas* has not said which signal, at which scope, or how far ahead of what.

  • Your growth is flat, so days of headroom is effectively infinite. Is the ceiling safe?
    No. The measure assumes growth stays smooth, and the events that consume a ceiling fastest are not growth at all: a failover, a recovery, a scale-up or a launch creates a large block at once. Alongside the trend, compare the granted ceiling against what your largest single planned action would consume.
  • Why alert per account and per region rather than on a single estate-wide figure?
    Because the ceiling is enforced per scope, so an aggregate hides the one scope that is about to refuse you. An estate at sixty percent overall can contain a region at ninety-eight percent. Compute the ratio where the ceiling lives, then aggregate for reporting only.
  • A raise you filed is still pending when the demand arrives. What now?
    Treat it as a capacity decision rather than a paperwork one: serve what the current ceiling allows, shed or defer the rest, and if the workload can run in another scope that already has headroom, place it there. The pending request is not a lever you can pull faster, so plan without it.

saying these in an interview costs you the question

  • Plans to raise the ceiling during the incident itself
  • Alerts on absolute usage instead of remaining headroom
  • Tracks only the accounts and regions used daily
  • Assumes a granted raise also reserves capacity
  • Averages headroom across regions into one figure
  • Leaves quota raises as nobody's backlog item