Your retail service hits its yearly peak on Black Friday, eight weeks away, and the forecast is roughly 4x your normal daily peak. Which parts of that capacity have lead times you cannot compress, and how would you sequence the eight weeks?
answer
- what has a queue in front of it?
- schedule backwards from the date
- quota is permission, not capacity
- stateful work is the critical path
- pre-scale; do not scale on the day
basics
~20 sCompute is usually the easy part. The long poles are cloud quota increases and capacity reservations for specific instance types and zones, stateful work like resharding or index builds, and third-party rate limits. Start those first and leave the final weeks for verification and a freeze.
solid answer
~50 sI start by turning the 4x forecast into resource requirements per tier, then ask which of them has a lead time measured in weeks rather than minutes. Typically: per-region, per-instance-family quota increases, which are a request that can take days and is not guaranteed; capacity reservations, because quota permits a launch but does not reserve physical capacity in a zone for a scarce shape; any stateful change — resharding, adding replicas, building an index, growing a cache tier — which is slow and risky; and third-party ceilings such as a payment processor's or SMS provider's rate limits, which need advance notice. Those go in weeks one and two. Weeks three to five are the stateful work. Weeks six and seven are verification against the target load and rehearsing the degrade plan. The last week is a change freeze with the capacity already scaled up and held, not scheduled to arrive on the day.
go deeper
Be ready to name a few things that take real time to obtain — quota increases, reserved capacity, database changes — and to say that a known peak is provisioned in advance rather than during the event.
Explain how a demand forecast becomes per-tier resource requirements, and why quota, scarce instance shapes and stateful work sit on the critical path while stateless replicas do not.
Lay out a credible backwards schedule with verification and a freeze in it, and describe what you do when an input arrives late or the shape you planned on is unavailable.
Own the org-level version: who tracks lead-time items across teams, how aggregate regional demand is reconciled with quota, and how much pre-provisioned cost the business is willing to carry against an uncertain forecast.
## The constraint is lead time, not money For a known annual peak, the budget conversation is usually the simple one. What kills teams is discovering in week seven that something they assumed was instantaneous takes three weeks. Capacity planning for an event is therefore a critical-path exercise: identify every input, attach a lead time to each, and schedule backwards from the date. ## The long poles **Quota.** Cloud accounts are limited per region and usually per instance family. Raising a limit is a request that a human or a system approves; it can take days, it can be partially granted, and it can be refused. Quota is also easy to forget for things that are not compute: addresses, load balancer rules, network throughput, managed-service instance counts, function concurrency. **Physical capacity for a shape.** Quota permits a launch; it does not reserve anything. A scarce instance type in a specific zone can simply be unavailable when you ask for two hundred of them at 9am on the busiest shopping day of the year. Capacity reservations exist precisely for this — you pay to hold the capacity ahead of time. Placing them requires knowing your instance family and zone weeks in advance, and qualifying a second family as a fallback is cheap insurance. **Stateful work.** Adding stateless replicas is fast. Resharding a database, adding and seeding read replicas, building a large index, growing a cache tier so its working set fits, expanding partitions on a log — these take days to weeks, carry real risk, and cannot be rushed into the week before a freeze. If the plan needs any of them, they are the critical path. **Third parties.** Payment processors, SMS and email providers, fraud-scoring APIs, identity providers and internal platform teams all enforce their own limits. Most will raise them for a known event, with notice. None will do it at 200% load on the day. **Commercial and legal.** Committed-use discounts, licence tiers priced per core or per node, and contract renegotiation all move on business calendars, not engineering ones. **People.** Staffing the rotation for the event, briefing everyone, and making sure the runbooks match reality are also lead-time items. ## Sequencing eight weeks - **Weeks 8-7 — forecast and translate.** Convert the 4x number into per-tier resource requirements: application cores, database queries per second and connections, cache working set, queue throughput, third-party call volume, egress. Identify which ceiling binds first. - **Week 7 — file everything with a queue in front of it.** Quota increases, capacity reservations across more than one zone and ideally more than one instance family, notice to every third party, and any commercial change. - **Weeks 6-4 — stateful work.** Resharding, replicas, index builds, cache growth. Land these early enough that a rollback is still possible. - **Weeks 4-3 — verify at target.** Drive the system at the planned load in a production-like environment and fix what breaks. Almost always the finding is not "we need more application servers"; it is a connection limit, a lock, an unindexed query or a downstream ceiling. - **Weeks 3-2 — rehearse the failure path.** Confirm the degradation and shedding controls work and that whoever is on call knows how to use them, and check that rollback for anything shipped in the window still works. - **Week 1 — freeze and pre-scale.** Stop risky change so the tested configuration is the one that runs, and bring the capacity up in advance rather than depending on it arriving during the surge. ## Pre-scale rather than scale on the day Even with elasticity, reactive scaling reacts: the metric window, the provisioning call, boot, dependency registration and warm-up all cost minutes, and the arrival curve for a promotion or a doors-open time is far steeper than that. Before a known peak, raise the floor — minimum fleet sizes, warm pools, pre-warmed serverless concurrency (AWS Lambda calls this provisioned concurrency), primed caches and established connection pools — so elasticity handles the surprise on top of a fleet that is already large enough for the plan. ## Have a plan for being wrong The forecast is an estimate. Decide in advance what you shed or degrade if demand exceeds the plan, who is allowed to pull that lever, and what the customer sees. A capacity plan without a defined behaviour above the plan is a plan that ends in an unmanaged failure. ## Afterwards Unwind on a schedule chosen in advance, and not before the tail — returns, support load and delayed fulfilment run for days after the sale. Then record the actual peak against the forecast so next year's number starts from data.
- Mid-window you discover the instance family you planned on is unavailable in one of your zones. What now?Fall back to the alternative family you qualified earlier, spread the reservation across more zones, or accept a different price point — and this is exactly why reservations go in week seven, not week one. If no fallback was qualified, benchmark one immediately, because cost per request differs by shape and the whole plan is denominated in that number.
- How does a change freeze interact with the capacity work?Every risky capacity change must land before the freeze starts, because the freeze is what makes the tested configuration the configuration that actually runs. The freeze should still allow the reversible operational levers — scaling up, flipping a degradation switch, rolling back — otherwise it removes your ability to respond during the event it was meant to protect.
- When do you give the capacity back?On a schedule decided in advance, and after the tail rather than the next morning — returns, support contacts and delayed fulfilment keep load elevated for days. Reservations and commitments have their own terms, so the exit conditions are part of the original purchase decision, not an afterthought.
saying these in an interview costs you the question
- Assuming you can simply scale up on the day
- Believing quota increases are instant and always granted
- Forgetting third-party and vendor rate limits
- Leaving resharding or index builds to the final week
- Provisioning application servers but not database connections