skip to content

A marketing launch will multiply your AWS workload's traffic next month. Using Service Quotas and Trusted Advisor, how do you make sure an AWS account limit is not the thing that fails first?

level: seniorimportance: must knowfreq 50%

answer

  1. Failure looks like throttling, not saturation
  2. Values differ per account and per Region
  3. Some limits will never be raised
  4. Lead time on increases is days
  5. Alarm on usage before launch day

basics

~20 s

Enumerate the quotas each service in the path consumes, read the applied values per account and Region in Service Quotas, and compare them against projected peak. Request increases on adjustable quotas weeks ahead, design around non-adjustable ones, and alarm on quota utilisation before launch day.

solid answer

~60 s

Start by listing every service on the request path and what quota each one consumes at peak — concurrent Lambda executions, running On-Demand vCPUs per instance family, API request rate at the front door, connections through a NAT gateway, free IP addresses in the subnets. Then read the *applied* values in Service Quotas, which are per account and per Region, so a workload spanning two Regions has two sets of numbers. Compare against projected peak with headroom, not the average. Adjustable quotas get an increase request now, because approval is a human process that can take days and some are routed into a support case; non-adjustable quotas are a design constraint you architect around — shard across accounts or Regions. Then make it observable: Service Quotas publishes usage into CloudWatch under the `AWS/Usage` namespace, and Trusted Advisor's limit checks give a second view, so you alarm at a fraction of the quota rather than discovering it at peak. Finally, load test in the account and Region you will actually launch in.

code

bash · 13 lines
bash
aws service-quotas list-service-quotas \
  --service-code lambda \
  --query "Quotas[?Adjustable].[QuotaName,QuotaCode,Value]" \
  --output table

aws service-quotas get-service-quota \
  --service-code ec2 \
  --quota-code L-1216C47A

aws service-quotas request-service-quota-increase \
  --service-code ec2 \
  --quota-code L-1216C47A \
  --desired-value 512

go deeper

for a junior

Know that AWS accounts have per-Region quotas, that hitting one shows up as throttling or a failure to launch resources, and that Service Quotas is where you look the values up.

for a middle

Be able to explain adjustable versus non-adjustable quotas, that applied values are per account and per Region, and that increases go through a review with real lead time.

for a senior

Demonstrate a readiness process: enumerate quotas along the request path, compare applied values to projected peak, request increases early, alarm on usage, and rehearse in the launch account and Region.

for a principal

Own quota management as an estate-wide concern — request templates for new accounts, a standing inventory of the limits your platform presses on, and a position on when sharding across accounts or Regions is the right answer to a hard limit.

## Why quotas are the classic launch failure A quota failure does not look like a capacity failure. Autoscaling quietly stops adding instances, function invocations get throttled, API calls start returning throttling errors, and the dashboards show a system that is *not* saturated. Everything looks healthy except what your customers are experiencing. That is why quota readiness is a distinct pre-launch activity rather than something a load test alone catches — unless you load test at real scale in the real Region. ## Step one: enumerate what the path consumes Walk the request path and, for each service, ask what unit it charges against a limit. Compute concurrency. Instance vCPUs by family. Load balancer targets and rules. Connections through a NAT gateway. Request rates at the API tier. Available IP addresses in each subnet — an underprovisioned subnet CIDR is effectively a quota, and it is not adjustable at all once the subnet exists. Also count the things that scale with success but sit off the request path: log ingestion, metric cardinality, queue and topic counts, tables and indexes. ## Step two: read the real numbers Service Quotas is the authoritative source. Two properties matter and both are routinely forgotten: - Quotas are **per account and per Region**. The us-east-1 value tells you nothing about eu-west-1. A multi-Region launch needs the exercise repeated per Region, and a multi-account estate per account. - Every quota has an **applied value** that may differ from the AWS default because someone raised it two years ago, and an **adjustable** flag saying whether it can move at all. Trusted Advisor's service-limit checks are a useful second view, but they cover a subset of quotas and refresh on a lag. Use them as a warning surface, not as the inventory. ## Step three: adjustable versus not For **adjustable** quotas, request the increase early through Service Quotas. The request is reviewed — sometimes automatically, sometimes by a human, sometimes routed into a support case — and lead time is measured in days, not minutes. Requesting the day before launch is how teams discover this. Ask for what you need at peak plus headroom, and be ready to explain the workload; a request for a hundredfold increase with no context attracts questions. For **non-adjustable** quotas, no amount of asking helps. These are architectural constraints: you shard across accounts, spread across Regions, or change the design so the limit is not on the critical path. Recognising early that a limit is hard is worth more than a week of negotiation. In an organisation, quota request templates let you pre-apply increases to newly created accounts, which stops every new team rediscovering the same three limits. ## Step four: make quota usage observable Service Quotas publishes usage metrics to CloudWatch in the `AWS/Usage` namespace, with the metric name `ResourceCount` and dimensions identifying the service and resource. You can alarm on usage relative to the applied quota, and the Service Quotas console will create that alarm for you for supported quotas. Trusted Advisor additionally publishes a `ServiceLimitUsage` metric in the `AWS/TrustedAdvisor` namespace. The rule is simple: anything the readiness review flagged as tight gets an alarm at a threshold well below the limit, wired to a channel a human reads. Discovering a quota at 100% during the launch is a failure of preparation, not of AWS. ## Step five: rehearse Load test in the Region and the account you will launch in, at the concurrency you expect, and watch specifically for throttling errors rather than only latency and error rate. Throttling usually surfaces as a distinct error code, and the number of teams whose load test ran in a pre-production account carrying different applied quotas is large. ## What a strong answer adds Two things. First, that headroom is not free of consequences: a very large On-Demand footprint concentrated in one Availability Zone can hit insufficient-capacity errors that have nothing to do with your account limits, so quota headroom and capacity availability are separate risks. Second, the after-action habit — record the applied values you actually needed, so the next launch starts from evidence instead of a blank page.

  • How does a non-adjustable quota change your plan compared with an adjustable one?
    An adjustable quota is a scheduling problem — request early, allow days for approval, verify the applied value afterwards. A non-adjustable quota is a design constraint: you either shard the workload across accounts or Regions so no single scope hits it, or you change the architecture so the limited resource is off the critical path. No support case will move it.
  • Your load test passed in staging but the launch throttled in production. What is the likely explanation?
    Different applied quotas. Quotas are per account and per Region, and a staging account often carries increases someone requested years ago, or the test ran in a different Region. Load tests must run in the account and Region you will launch in, or at minimum you must diff the applied quota values between the two before trusting the result.
  • How would you know you were approaching a quota before customers did?
    Alarm on the Service Quotas usage metrics CloudWatch publishes in the `AWS/Usage` namespace, at a threshold well below the applied quota, for every limit the readiness review flagged as tight. Trusted Advisor's limit checks and its `ServiceLimitUsage` metric give a second, coarser signal. The point is that the alarm exists before launch day, not that it is perfect.
  • Why is subnet sizing worth treating as a quota question?
    Because free IP addresses behave exactly like a limit: when a subnet runs out, tasks or instances simply fail to launch, and an existing subnet's CIDR cannot be resized. It is effectively a non-adjustable quota you set yourself at design time, and it bites hardest on workloads where every task or pod consumes an address.

saying these in an interview costs you the question

  • Assumes quota increases are instant and automatic
  • Reads quotas in one Region and assumes the account matches
  • Treats every quota as raisable on request
  • Relies on Trusted Advisor as the complete quota inventory
  • Load tests in an account with different applied quotas

context