A platform lead argues that because every service now scales automatically with load, the company can drop its capacity planning process entirely. You are accountable for both reliability and infrastructure spend — how do you answer?
answer
- scaling allocates; it does not create
- the pool has a ceiling
- reacts in minutes, spikes in seconds
- commitments are a bet on a forecast
- who owns the sum of every maximum?
basics
~20 sAutomatic scaling allocates capacity that already exists; it does not create it. Account and region quotas, physical zone capacity, stateful tiers and vendor limits all set ceilings, scale-up takes minutes, and buying multi-year commitments requires a demand forecast in the first place.
solid answer
~50 sI would concede the part he is right about: nobody should be hand-counting instances per service any more, and per-service sizing is better delegated to a control loop than to a spreadsheet. Then I would name what the loop cannot do. It cannot raise an account or region quota, conjure a scarce instance shape into a zone, reshard a database, or lift a payment provider's rate limit — those are the real ceilings and they all have lead times. It reacts in minutes, so anything with a step-shaped arrival still needs a floor set from a forecast. It says nothing about money: committing to one- or three-year discounts is a bet on a demand floor, and without a forecast you either pay list price on steady baseline or over-commit and eat the sunk cost. And somebody has to own the aggregate — the sum of every team's configured maximum has to physically exist in the region. So the process shrinks and changes shape; it does not disappear.
go deeper
Know that automatic scaling works within limits somebody configured, and that quotas and databases can run out even when the scaling maximum is set high.
Be able to name the concrete ceilings — account and region quota, instance availability in a zone, stateful tiers, third-party rate limits — and explain why reaction time makes a pre-set floor necessary before known events.
Demonstrate the operational consequence: pre-scaling before events, verifying that scaling maximums are actually achievable, and recognizing that the stateful tier is usually the binding constraint.
Own the position end to end — what the planning process is reduced to, who owns the aggregate regional number, and how much of the fleet is committed versus on-demand versus preemptible as a forecast-driven financial bet.
## What the loop actually does Automatic scaling is a feedback controller. It observes a signal, compares it to a target, and adjusts a count within a configured range, drawing from a pool of capacity that someone already arranged to exist. Every part of that sentence has a limit in it: the signal lags, the range has a maximum, and the pool is finite. Capacity planning is a different activity. It decides how large the pool should be, where it should physically live, what it should be bought under, and which of its inputs need to be ordered weeks ahead. Elasticity makes planning less granular. It does not make it unnecessary. ## Ceilings the controller cannot raise - **Account and region quota.** Limits on cores per instance family, addresses, load-balancer objects, managed-service instances, function concurrency. A scaling maximum set above the quota is fiction, and the failure surfaces as a quota error during a surge. - **Physical capacity for a shape.** Quota is permission; it is not a reservation. A scarce instance type in a specific zone can be unobtainable at the exact moment demand arrives. - **Stateful tiers.** Databases, shard counts, log partitions and cache tiers do not scale on the same timescale as stateless replicas, and frequently not automatically at all. In most architectures they are the true capacity limit, so scaling the stateless tier into them makes the incident arrive faster. - **Third-party and internal platform limits.** Payment processors, messaging providers, identity services and shared internal platforms enforce their own rates. Scaling past them converts your capacity into their throttling. - **Licences and contracts.** Per-core or per-node pricing tiers move on business calendars. ## Reaction time versus arrival shape A control loop needs a metric window, then a provisioning call, then boot, registration and warm-up. That is minutes for most stacks. Demand from a push notification, a doors-open sale or an upstream failover arrives in seconds. The mitigation is not a faster loop, it is a **floor derived from a forecast**: minimum fleet sizes raised ahead of known events, warm pools, pre-warmed serverless concurrency, primed caches and established connection pools. Elasticity then handles the surprise on top of a fleet that was already correct for the plan. ## The money argument is the stronger one Even granting perfect elasticity, the purchasing decision is unavoidable. Discounted long-term commitments are priced against a term — you are betting that a floor of demand exists for one or three years. That bet requires a forecast to make at all: - Commit too little and you pay on-demand rates on a steady baseline that was never going away. - Commit too much and a demand miss becomes sunk cost you cannot scale away from. The usual shape is to commit to the confident floor, run the volatile middle on-demand, and put interruptible work on preemptible or spot capacity — but every one of those boundaries is a number that comes out of a forecast. "We autoscale" is not an input to that decision. ## Somebody must own the aggregate Each team can size its own service correctly and the company can still fail, because the sum of every team's configured maximum has to physically exist in the region on the worst day. No single team can see that total. That is the irreducible central function: collect per-service forecasts, convert them to resources, aggregate by region and by instance family, compare against quota and reservations, and track the lead-time items. It is also where launches get caught — the campaign that four teams are unknowingly serving at once. ## What I would actually propose Replace the heavyweight process with a light one that keeps only the parts elasticity cannot do: 1. A short forecast per significant service, refreshed quarterly, expressed in resource units. 2. Aggregation per region against quota and reservations, owned by the platform or SRE group. 3. A calendar of demand events — launches, campaigns, seasonal peaks — with lead-time items tracked to a date. 4. A purchasing position revisited each cycle: what fraction of the fleet is committed, on-demand, or preemptible. 5. Explicit ownership of the stateful tiers' capacity, because that is where the real ceiling usually is. ## How you know the process was dropped too early The signature is distinctive: capacity incidents that show up as quota errors, unavailable instance types or exhausted database connections rather than as high utilization; scaling maximums nobody can justify; and the platform team learning about a launch from the traffic graph. If those are appearing, the planning function was not eliminated — it was just left unowned.
- What fraction of the fleet would you cover with long-term commitments?Roughly the floor the forecast says will exist for the whole term, with the volatile portion on-demand and interruptible work on preemptible capacity. Over-committing turns a demand miss into sunk cost; under-committing pays list price on baseline that was never going away. Provider terms differ in flexibility, so the mix is revisited each planning cycle rather than set once.
- Which group should own the aggregate regional capacity number?A group with cross-team visibility — platform or SRE. No individual service team can see that everyone's configured maximum sums past the region's quota or the reserved capacity available. Teams own their own forecasts and their stateful tiers; the central group owns aggregation, the long-lead items, and the event calendar that catches simultaneous launches.
- What early signal tells you the planning function has been dropped rather than automated?Capacity incidents that present as quota errors, unavailable instance shapes or exhausted database connections instead of as high utilization; scaling maximums nobody can explain; and the platform team discovering a launch from the traffic graph. Those all indicate the ceilings and lead-time items stopped having an owner.
saying these in an interview costs you the question
- Treating elastic capacity as effectively unlimited
- Ignoring per-region quota and physical zone capacity
- Assuming stateful tiers scale as fast as stateless ones
- Buying multi-year commitments with no demand forecast
- Leaving the aggregate cross-team number unowned