skip to content

Elasticity is meant to protect responsiveness, yet scaling out mid on-sale worsened the checkout request's 99th-percentile latency - how would you settle that trade-off?

level: principalimportance: nice to knowfreq 33%

answer

  1. one goal, three instruments
  2. write the bound down first
  3. new capacity is not immediately useful
  4. the cost peaks when load rises fastest
  5. scheduled load is forecast, not reacted to

basics

~20 s

Rank the properties before arguing: responsiveness is the goal and elasticity is an instrument judged only by whether the bound holds. State the bound, measure the scale event as a cost against it, and for a scheduled spike provision ahead rather than react.

solid answer

~60 s

Start by refusing to treat the two as equals. Responsiveness is the property with a user-facing definition, and elasticity exists to keep it true as load moves; if a scale event breaks the bound, the instrument is spending the goal it was bought to protect. So I would write the bound down as the acceptance criterion - this percentile of this request, at this arrival rate, including during a single dependency failure - and then treat the reconfiguration window as a measured cost, because new capacity is not useful the moment it exists: caches are cold, connections must be established against shared components, and work in flight is redistributed, all of which lands hardest while load is rising fastest. Then I would ask what the added instances contend on, since a shared serialized component turns scale-out into a contention problem that shows up first in the tail. And because an on-sale is on the calendar, I would provision ahead of it and leave the elastic mechanism for the unpredictable residue and the scale-down afterwards.

go deeper

for a junior

Take away the ranking rather than the tactics: scaling is something a system does in order to keep answering quickly, so if it stops answering quickly, the scaling has not helped no matter how many instances are running.

for a middle

Explain why capacity added mid-spike is not immediately useful - cold local state, new connections against shared components, redistribution of work in flight - and why those costs land in the tail rather than the mean.

for a senior

Measure the scale event itself: run it under load off-peak and record the percentile through the reconfiguration window, then look for the shared component the new instances contend on when the tail worsens.

for a principal

Own the ranking and the bound. Decide what the organisation defends, require each instrument to ship with the evidence it did not spend that bound, and make the same call for the headroom that isolation reserves.

## Rank the four before you argue about them The argument only looks hard if elasticity and responsiveness are treated as two goods to be balanced. They are not peers. Responsiveness is the goal; elasticity is one of two means that keep it true - in this case, as load moves. That ranking settles the question of what evidence decides: an elastic mechanism that makes the checkout request's high percentile worse has failed on its own terms, whatever it did to utilisation or cost. So the first move is to write the goal down as something that can be checked: - **which request** - the checkout request, end to end as a user experiences it; - **which percentile and bound** - the 99th, under some stated figure; - **at which arrival rate, over which window** - the on-sale peak, for its duration; - **under which failure** - the bound must survive at least one dependency failing, since resilience is the other means. Everything after this is an instrument measured against that sentence. ## Why added capacity can cost latency while it is being added New capacity is not useful the instant it exists, and the gap is where the tail comes from: - **Cold local state.** A new instance starts with nothing cached or precomputed, so its early requests are its slowest, and they are dealt to real users. - **Connection and session establishment.** Every new instance opens connections to whatever is shared behind it, and a burst of them arrives at the shared component exactly when it is already busiest. - **Redistribution.** Whatever spreads work across instances has to change its mind, and work in flight during the change is the work that waits. - **Coordination cost.** More participants means more of whatever they must agree on, and a shared serialized component gets more contended rather than less. All four are worst when load is rising fastest, which is exactly when a reactive mechanism decides to act. That is the trap: the instrument fires at the moment its cost peaks. ## A scheduled spike is not the case reactive scaling is for An on-sale is on the calendar. The load is known in advance, within a factor that historical curves supply, which makes reacting to it the wrong instrument: 1. **Provision to the known peak before the window opens**, so the reconfiguration cost is paid while nobody is waiting. 2. **Leave the elastic mechanism for the residue** - the error in the forecast, and the traffic that arrives for reasons nobody predicted. 3. **Use it for the way down.** Releasing capacity afterwards without the bound regressing is the half of elasticity that gets skipped, and it is the half that shows nothing in the path silently depends on the extra capacity staying. | Load shape | Instrument that fits | What to measure | |---|---|---| | Known, scheduled peak | provision ahead of the window | the bound through the whole window, not the average | | Unpredictable growth | reactive scale-out, with the reconfiguration window budgeted | the bound during the scale event, separately | | Decay after a peak | scale-down | the bound after release, and whether anything regressed | ## Find the shared serialized component before adding instances If the added instances all contend on one thing - a single coordinating component, one shared store, one lock-shaped resource - then scale-out converts a capacity problem into a contention problem, and the tail is where contention shows up first. Elasticity, as a property, includes the absence of such a point in the path: a system with one is not elastic, it is merely deployable in multiples. The diagnostic is straightforward: if the high percentile worsens as instances are added while the mean barely moves, the added instances are queueing for something they share. ## Making the decision, and making it stick 1. **Name the bound and who may spend it.** Capacity, cost and utilisation are all negotiable against each other; the bound is the thing the organisation defends. 2. **Require every means to ship with its measurement.** A scale-out policy arrives with the percentile measured through a scale event, run off-peak, or it arrives unproven. 3. **Budget the reconfiguration window explicitly.** If the window cannot be made shorter than the spike is steep, the mechanism must fire earlier - on a forecast - rather than on the load itself. 4. **Accept the symmetric cost on the other means.** Isolation, which buys containment, reserves capacity that cannot be pooled, so resilience is paid for in headroom. That is the same kind of decision, made once, at the property level. The answer an interviewer is listening for is not a scaling policy. It is the ranking: one property is the goal, the other three are instruments, and when an instrument costs the goal, the instrument is what changes.

  • The mean barely moved while the 99th percentile worsened as instances were added. What does that pattern suggest?
    That the added instances are queueing for something they share. Contention appears in the tail long before it appears in the mean, so a worsening percentile against a flat mean points at a shared serialized component in the path rather than at a shortage of capacity.
  • Does this reasoning make reactive scaling a mistake in general?
    No. It is the right instrument for load nobody forecast and for releasing capacity afterwards. The point is that it has a cost concentrated in the reconfiguration window, so for load that is known in advance you pay that cost before the window opens instead of during it.
  • How does the same argument apply to resilience rather than elasticity?
    Identically. Isolation buys containment by reserving capacity that cannot be pooled, so resilience is paid for in headroom and, at the margin, in cost. It is judged the same way: does the response-time bound hold during an injected failure? If containment never has to work, it was headroom bought for nothing.

saying these in an interview costs you the question

  • Treating the four properties as equally weighted goals
  • Assuming added capacity helps the moment it starts
  • Reacting to a spike that was on the calendar
  • Reading a worsening tail as measurement noise
  • Adding instances without finding the shared serialized component