skip to content

What separates a soft quota you can ask to have raised from a hard limit, and what does hitting each cost?

level: juniorimportance: must knowfreq 68%

answer

  1. two ceilings, two rescue plans
  2. one is policy, one is design
  3. first question: can it be raised?
  4. raisable costs days of lead time
  5. fixed costs a restructure, not a request

basics

~20 s

A soft quota is a provider ceiling an increase request can raise, so hitting it costs lead time. A hard limit is fixed by the platform's design, and the only way past it is an architecture change.

solid answer

~50 s

Both are ceilings the platform enforces on the creation path, so the symptom is the same: a provisioning call is refused and nothing gets built. The difference is what you can do about it. A **soft quota** is a number the provider chose as a starting value, counted per account and usually per region, and there is a documented way to have it raised — so the real cost of meeting one is *time*, typically days spent waiting on the provider's queue. A **hard limit** falls out of how the service itself is built, and no request will move it; getting past one means changing the shape of the design, such as spreading the workload across more than one account or region. Classify every ceiling your design leans on, because one is a scheduling problem and the other is a design constraint.

go deeper

for a junior

Recall the two words and the one difference: a soft quota can be raised on request, a hard limit cannot. Say which of the two you would look up first when a create call is refused.

for a middle

Explain the mechanics: the ceiling is enforced when the resource is created, the raise is granted for a scope rather than globally, and it may be granted only in part. Put a number on the lead time in days.

for a senior

Show the operating habit. Name the handful of ceilings your design leans on, say who owns filing the raises, and describe what you would do differently the moment a ceiling turns out to be fixed rather than raisable.

for a principal

Frame it as a scheduling risk against a design risk. Decide which ceilings the organisation tracks centrally, and argue where a fixed ceiling should force a structural choice early rather than be absorbed by growing headroom requests.

## Two ceilings that produce the same error Every large platform caps how much of a given thing one customer may have: how many machines of a family may run at once, how many private networks may exist in a region, how many rules one filter may hold. The cap is enforced on the **creation path** — the management API refuses the request and nothing is built — so the first symptom is a failed provisioning call, not a failing workload. Anything already running keeps serving. What the refusal rarely makes obvious is **which kind of ceiling you just met**, and that is the one fact that decides what you do next. There are two kinds. A **soft quota** is a number the provider chose and is willing to change. A **hard limit** is a number that falls out of how the service is built. ## A soft quota is a default someone chose for you A soft quota is a starting value applied to your account, counted separately in each region for most resources. It exists for two reasons at once: it protects the platform from one customer's runaway automation, and it protects the customer from a mistake that would otherwise arrive as a bill. Because it is a policy number rather than a structural one, there is a documented way to change it — on some platforms a self-service form, on others a support case, with the review either automatic or handled by a person who asks what you intend to run. The consequences that matter in practice: - The raise is granted **for a scope**, typically one account in one region. It does not spread to your other accounts or your other regions. - A raise can be granted **partially**: you ask for a large number, receive a smaller one, and are invited to come back with usage evidence. - The real cost of meeting a soft quota is **time**. The engineering work is filling in a form; the delay is the provider's queue, and it is usually measured in days. - Because the cost is time, the mitigation is **calendar work** — request before you need it, rather than when the creation call fails. - A granted raise is permission to ask for more. It is not a reservation of anything. ## A hard limit is a property of the design A hard limit is fixed by how the service is built: an addressing scheme with a finite size, a partitioning scheme with a fixed count, a maximum size for one object, a maximum number of things attachable to one other thing. Providers normally publish these alongside the raisable ceilings and mark them differently. There is no request to file, and asking produces a polite no rather than a delay. The cost is therefore not time but **rework**. Getting past a hard limit means changing the shape of the design — splitting across more than one account or region, sharding what you were treating as a single object, or choosing a construct that is not bounded the same way. That is engineering measured in weeks, which is why meeting a hard limit late is far more expensive than meeting a soft quota late. ## The comparison at a glance | | Soft quota | Hard limit | |---|---|---| | Where the number comes from | a provider default applied to your account | how the service itself is built | | Can it be raised | yes, through a documented request | no, and asking does not help | | Cost of meeting it | lead time, usually days | redesign, usually weeks | | Correct response | request early, track headroom | treat the number as a design input | | What it protects | the platform, and your bill | the feasibility of the service | ## What to do before you meet either 1. List the ceilings your design actually leans on — not every published number, only the handful your growth curve will reach. 2. Classify each as raisable or fixed. That is one lookup per ceiling, and everything downstream depends on it. 3. For the raisable ones, find out what a raise takes and put the request in the backlog with an owner, ahead of the event that needs it. 4. For the fixed ones, write the number into the design as a constraint the way you would a maximum payload size, so you approach it deliberately instead of discovering it. ## Why this is asked The two failures have completely different rescue plans, and teams routinely apply the wrong one. Answering *file a request* for a hard limit commits you to waiting for something that will never arrive; answering *redesign* for a soft quota spends weeks of engineering on a problem a form would have solved. The follow-up is nearly always about time: how long does a raise take, and what were you doing to make sure you asked for it before the launch rather than during it.

  • If a raise is usually granted, why not simply wait until the creation call is refused and file it then?
    Because the request lands on the provider's queue, not yours. The refusal arrives at the worst moment — a launch or a failover — and the only lever left then takes days, during which nothing you can do speeds it up. Filing ahead of the event converts a blocking incident into a routine backlog item.
  • A raise you asked for was granted at a lower number than you requested. What does that tell you?
    That the ceiling is discretionary rather than an entitlement: the provider reviews what you intend to run and grants against that. Treat the smaller number as a checkpoint — run at it, keep the usage evidence, and re-ask. It also means you cannot assume a single request will carry you to your target.
  • Does a granted raise guarantee the resource will be created?
    No. Raising a soft quota removes your ceiling as the blocker; it does not reserve anything on the provider's side. The refusal you see afterwards, if any, is a different failure with a different message and a different remedy, so read the error rather than assuming the raise did not apply.

A card with a spending limit the bank will raise if you call is not the same as a card that physically cannot store more than sixteen digits. One costs a phone call and a wait; the other means you need a different card.

saying these in an interview costs you the question

  • Says every ceiling can be raised if you ask nicely
  • Treats a raise as effective the moment the request is submitted
  • Answers a hard limit with a support request and waits
  • Assumes a raise applies to every account and region at once
  • Plans a redesign around a ceiling a request would have raised