On the managed ladder's three rungs, traffic doubles for one service — how does adding capacity differ at each rung?
answer
- same verb, three shapes
- object, target, consequence
- you build; platform reconciles; demand decides
- the ceiling is what stays yours
- mechanics handed over, scalability not
basics
~20 sCapacity changes from an object you build, to a target you declare, to a quantity you only meet on the bill. The rung decides who performs the scaling, never whether the service must tolerate copies coming and going.
solid answer
~50 sOn the rented-machine rung capacity is a thing you own: you choose a bigger machine or build more of them, register them with whatever sits in front, and retire them yourself, and the lead time is provisioning plus boot plus warm-up — your lead time. On the managed-runtime rung you declare a target instead: a copy count, or a range plus a signal such as utilisation or queue depth, and the platform reconciles toward it, placing, health-checking, replacing and retiring copies on your behalf. On the whole-capability rung capacity is often not an object you name at all — it follows demand and shows up afterwards as a quantity you are charged for, and what you own shrinks to a ceiling. What never moves up the ladder is the requirement that copies can appear and disappear safely.
go deeper
Know that the higher the rung, the less you do by hand: first you build machines, then you declare a target, then you do nothing and read the quantity you were charged for.
Explain the reconciliation loop on the middle rung — a declared target, a watched signal, the platform placing and retiring copies — and say which lead time is still yours.
Bring the failure with you: a policy that reacted after the spike, a warm-up nobody measured, a ceiling met at the worst moment. Say which signal you would drive the target from.
Decide where capacity risk should sit across the estate: a target teams own and can get wrong, or a provider-driven curve nobody can pre-warm and everyone can only cap.
## Same verb, three shapes "Scale" sounds like one operation, but the thing you actually do differs completely by rung. What changes is not how fast capacity arrives — that varies for other reasons — but **what kind of thing capacity is**: an object, a declaration, or a consequence. ## Rung one: capacity is an object you own On a rented machine, capacity is countable and you hold each one. - You either grow the machine (**right-sizing** it up, usually with a restart) or you build more machines. - Each new machine has to be created from an image, configured, started, warmed up, and registered with whatever distributes traffic to it. - When demand falls, nothing removes them but you, and a machine nobody retires goes on being charged for. - The lead time is yours end to end: provisioning, boot, application start, cache warm-up. This is the rung with the most control and the most homework. It is also the only rung on which you can hold extra capacity idle in advance simply because you expect a spike, without asking anyone. ## Rung two: capacity is a target you declare On a managed runtime you stop creating copies and start describing how many there should be: a fixed count, or a range together with a signal the platform watches — utilisation, request concurrency, queue depth. A reconciliation loop does the rest. - The platform places each copy on a host, health-checks it, retires unhealthy ones and replaces a host that dies. - Removal happens as readily as addition, which is where teams get hurt: copies are destroyed on the platform's schedule, not yours. - What remains yours is the **policy** — which signal, what thresholds, what floor and ceiling — and the consequences of getting it wrong. - The reaction time you own is the policy's, not the machine's: how long the signal takes to move, and how long a new copy takes to become useful. ## Rung three: capacity is a consequence of demand On a whole managed capability there is often nothing to declare. Capacity follows the load and appears afterwards as a quantity in the charge model. You usually cannot name an instance, cannot pre-warm one, and cannot hold a shape through a known spike. What you do still own is a **ceiling**: nearly every such tier lets you bound what it will scale to, because unbounded demand is as dangerous to you as too little capacity. ## Side by side | question | rented machine | managed runtime | whole capability | |---|---|---|---| | what you change | the machines | the target | usually nothing | | who adds copies | you | the platform | the provider | | lead time you own | provisioning, boot, warm-up | the policy's reaction time | none you control | | can you pre-warm | yes, freely | often, via a floor | usually not | | what a spike leaves behind | machines you must retire | copies inside your range | a quantity on the bill | ## What the hand-over does not do This is the part interviews are really testing. Handing over the mechanics of scaling does not make a service scalable. 1. **Copies must be interchangeable.** A request has to be servable by any of them. 2. **A copy must be safe to destroy mid-flight.** A session, a cache or a work item held only in local memory is lost when a copy goes, and the higher rungs destroy copies more often, not less. 3. **Warm-up is still warm-up.** If a new copy needs minutes before it is useful, no rung shortens that; the platform simply starts the clock without asking you. 4. **Demand is still yours.** No rung reduces how much capacity the workload needs; it only changes who arranges it. ## The failure to describe in an interview The honest scenario is not "we scaled up". It is a policy that reacted after the spike had passed, because the signal was slow and the warm-up was long, on a tier where nothing could be pre-warmed. Saying which signal you would drive the target from, and what floor you would hold to cover the warm-up, is what distinguishes someone who has operated this from someone who has read about it.
- What still has to be true of the service before any rung can add copies?Copies must be interchangeable: a request can land on any of them, and one can vanish mid-flight without taking anything irreplaceable with it. A service that keeps a session, a cache or a work item in local memory is not made safe by moving up a rung — the platform simply destroys it more often.
- Which capacity control survives on the top rung?The ceiling. Even where you cannot name or pre-warm an instance, a tier normally lets you bound the maximum it will scale to, because unbounded demand is as dangerous to you as too little capacity. Setting that bound is the one capacity decision the rung leaves you.
saying these in an interview costs you the question
- Thinks a higher rung makes a service with local state safe to replicate.
- Assumes the top rung removes any ceiling on capacity.
- Believes declaring a scaling target removes a new copy's warm-up time.
- Treats scaling as free once the platform performs it.
- Expects to pre-warm capacity on a rung that never exposes an instance.