A store's provisioned throughput floor was sized for a peak that ended a quarter ago - how do you separate dead capacity from headroom somebody holds on purpose?
answer
- you buy the floor, not the traffic
- the ratio looks the same either way
- ask for the event, not the average
- approaches to the floor over a cycle
- step down, and check the raise path
basics
~20 sUtilisation alone cannot tell them apart. A low ratio only proves the floor exceeds observed demand; deliberate headroom has a named reason and an event that consumes it. Get consumption across a full business cycle plus an owner's stated reason, then step the floor down rather than cutting it.
solid answer
~50 sYou are charged for the floor you provisioned, not the traffic you served, so a floor sized for an ended peak bills in full forever. But a low utilisation ratio is not evidence of waste on its own: deliberate headroom looks identical in a chart. The distinguishing facts are a **reason** and a **consuming event** - a failure-domain loss that shifts load onto what remains, a seasonal peak, a batch window - which show up as the floor being approached at predictable moments across a full business cycle. Flat consumption far below the floor, with no owner able to name the event, is stranded. Reduce it in steps rather than to the observed peak, check first how quickly the floor can be raised again, and be aware that if the usage was already bought under a term commitment, cutting consumption may not change the bill at all.
go deeper
Recall that a provisioned floor bills for the capacity you declared, not the traffic you served, so a floor set for an old peak keeps charging in full.
Explain why a utilisation ratio cannot distinguish stranded capacity from deliberate headroom, and what a peak-of-interval view over a full cycle shows that an average hides.
Demonstrate the safe reduction: ask the owner for the event the headroom covers, step the floor down across cycles, check how quickly it can be raised again, and watch throttling rather than latency.
Own the asymmetry - an unnecessary cut is found during an incident, an unnecessary floor is found in a bill - and insist that every floor carries a recorded reason so the next sweep inherits a decision.
## What the floor actually charges for Some resources are billed on a capacity you declare rather than the work you do: a throughput floor on a managed store, a reserved pool of connections, a fixed number of always-warm instances, a minimum node count. The bargain is predictable performance for a predictable price, and the charge runs on wall-clock time whatever the traffic. When the workload that justified the floor is redesigned or retired, the floor does not notice. It is the quietest form of orphaned spend, because unlike a detached volume the resource is genuinely in use - just at a fraction of what you are paying for. ## Why a utilisation ratio is not a verdict Here is the trap. Two situations produce the same chart: - **Stranded capacity.** The peak that justified the floor happened once and will not return; nobody has thought about the number since. - **Deliberate headroom.** Somebody sized the floor for an event that has not occurred yet - a failure domain being lost and its load landing on the survivors, a seasonal peak, a nightly batch, a launch. Both show consumption far below the floor for months. Cutting the first saves money; cutting the second removes the protection on the day it was bought for, and you find out during the incident. So the evidence has to be about the **event**, not the ratio. | Signal | Stranded | Deliberate headroom | |---|---|---| | Consumption shape over a cycle | Flat, far below the floor, no approaches | Approaches the floor at identifiable moments | | Reason on record | None; the number predates a redesign | Named, with the event it covers | | Owner's answer | "That was for the old pipeline" | "That is what we need when a zone is lost" | | What happens if you cut it | Nothing, ever | Nothing, until the event | This leaf's question is only whether **anybody ever chose** the number. How much headroom a service ought to hold, and the utilisation target it should run at, is a capacity-planning judgement with its own owner; the reconciliation's job is to establish that the number is a residue rather than a decision, and then hand it to whoever should own it. ## Getting the evidence 1. Pull **consumption against the floor across a full business cycle**, at a resolution fine enough to see short peaks. An hourly average can hide a burst that repeatedly touched the ceiling; a peak-of-interval view will not. 2. Look specifically for **approaches to the floor** - how often, and at what times. Regular approaches at month end or during a nightly window are the signature of real demand. 3. Ask the owner for the **event** the headroom covers, and check whether that event is still possible. "If we lose a failure domain the rest absorbs the load" is a reason; "we set it during the migration" is a residue. 4. Check whether the workload that set the number still exists at all. A floor sized for a pipeline that was redesigned a quarter ago is the clean case. ## Cutting it without creating an incident - **Step it down, do not jump to the observed peak.** Halve the gap, watch a cycle, halve again. The saving arrives slightly later and the risk falls a great deal. - **Check the raise path before the cut.** Raising some floors is immediate; on others it takes a rebuild, a migration or a support request, and that asymmetry - cheap to cut, slow to undo - should make you step more cautiously, not less. - **Watch the right signal during the step**, which is usually throttling or queueing against the floor rather than latency alone, since the first symptom of an undersized floor is requests being held back. - **Record the new number with its reason**, so the next sweep finds a decision rather than another residue. ## The case where cutting changes nothing One outcome surprises people: the consumption is reduced, the floor comes down, and the bill does not move. That happens when the usage was already paid for under a spend or usage commitment made earlier - the money is committed whether or not you consume it, so reducing consumption converts an over-provisioned resource into unused commitment rather than a saving. Sizing, valuing and unwinding such a commitment is a separate subject with its own trade-offs; the point for a sweep is simply to check whether the line you are about to reduce is already covered, before promising anyone a number.
- Why is a monthly average utilisation figure a poor input to this decision?Because the whole question is about peaks. An average flattens exactly the short approaches to the floor that distinguish real demand from residue, so a deliberately sized floor and a stranded one look identical. Use peak-of-interval consumption across a full cycle, and count how often the floor was approached.
- What makes stepping a floor down safer than cutting it to the observed peak?Each step is a smaller change with a cycle of evidence behind it, and it keeps margin for a demand shape you have not observed yet. It also matters because raising some floors is slow or needs a rebuild - cheap to cut and expensive to undo argues for several small moves rather than one large one.
saying these in an interview costs you the question
- Low utilisation proves the capacity is wasted.
- Cut the floor straight to last quarter's observed peak.
- Reducing provisioned capacity always reduces the bill.
- An average over the month is enough evidence.
- If nothing broke after the cut, the cut was right.