skip to content

A workload's copy count swings between four and twelve every few minutes - what causes that, and what damps it?

level: middleimportance: should knowfreq 46%

answer

  1. the loop is reacting to itself
  2. removal is instant, addition is slow
  3. high-water mark, not a cooldown
  4. grow fast, shrink slowly
  5. flat total load means it is self-induced

basics

~20 s

The loop is reacting to its own actions. Adding copies drives the per-copy signal below target, the loop removes them, the signal jumps back above target, and it adds again. A settling window, a tolerance band around the target, and slower shrinking than growth damp the oscillation.

solid answer

~40 s

Flapping is a closed loop oscillating: the signal the loop reads is changed by the action the loop takes, and the two are separated by a delay. Add copies, per-copy load falls, the recompute now asks for fewer, copies are removed, the same load lands on fewer copies, the signal jumps back over target. A noisy short-window signal and symmetric aggressiveness both amplify it. The standard damping is a **settling window**: before shrinking, take the *highest* recommendation over the last several minutes and only shrink to that. Pair it with a tolerance band so small deviations do nothing, and with asymmetric rates - grow quickly, shrink slowly. The cost is deliberate: you keep paid-for copies after load has already dropped.

code

pseudocode · 14 lines
pseudocode
every evaluationInterval:
    recommendation = recomputeDesiredCount()
    history.record(now, recommendation)

    if recommendation > currentCount:
        setDesiredCount(recommendation)          # grow immediately
    else:
        settled = max(history.since(now - settlingWindow))
        if settled < currentCount:
            setDesiredCount(settled)             # shrink only to the window's high-water mark
        # otherwise a higher recommendation inside the window blocks the shrink

# currentCount = 12, window holds 12, 12, 5, 4 -> max = 12 -> no shrink
# once the 12s age out: window holds 5, 4 -> max = 5 -> shrink to 5

go deeper

for a junior

Know that a scaling loop can argue with itself: the copies it adds lower the number it is watching, which tempts it to remove them again a minute later.

for a middle

Walk the oscillation cycle in order and explain the settling window as a high-water mark over recent recommendations, applied when shrinking but not when growing.

for a senior

Diagnose before tuning: compare total offered load against copy count to separate self-induced oscillation from genuine burstiness, and account for the cost of discarding warm capacity.

for a principal

Own the asymmetry as policy - the estate should be systematically eager to grow and reluctant to shrink, and the standing cost of that reluctance is a deliberate purchase of stability.

## Why a control loop oscillates This is not a bug in the arithmetic; it is the classic behaviour of a feedback loop whose own output moves its input, with lag in between. The cycle runs like this: 1. Load rises, the per-copy signal crosses its target, and the loop recomputes upward - say from 4 to 12. 2. The same total load now divides across 12 copies, so the per-copy signal drops well under target. 3. The next evaluation reads that low value and recomputes downward - back toward 4. 4. The load, unchanged, now divides across 4 copies again, and the signal jumps over target. 5. Back to step 1. Three things make it worse: - **A short or noisy measurement window.** The less smoothing, the more the loop chases samples rather than levels. - **Symmetry.** Growing and shrinking at the same speed means the loop is as eager to give capacity back as it was to take it. - **Asymmetric real cost.** Removal is nearly instant; addition costs placement, start-up, readiness and warm-up. So the fleet spends a meaningful share of its time *below* the count it actually needs, waiting for copies it just discarded to come back. ## Flapping is not merely wasteful - it can serve worse than a fixed count Each removal drains and terminates a copy that was carrying live work; each addition pays the full cold chain again. In a bad oscillation, a fleet can spend most of its time either short of capacity or mid-start, and a static count sized for the peak would have served the same load with fewer errors. That is the argument to make when a team proposes tuning the loop to be more aggressive after a slow scale-up incident: aggressiveness converts a lag problem into an oscillation problem. ## The damping controls | Control | What it does | What it costs | |---|---|---| | Settling window | Before shrinking, take the highest recommendation seen over the last N minutes and shrink only to that | Copies are held after load has genuinely dropped | | Tolerance band | Ignore deviations within a few percent of the target | The fleet sits slightly off target most of the time | | Asymmetric rates | Grow fast, shrink slowly - for example at most one copy removed per interval | Slower cost recovery after a peak | | Longer averaging window | Smooths the signal before the loop ever sees it | Slower reaction to a genuine rise | The settling window is the load-bearing one, and its shape matters: it is a **high-water mark over recent history**, not a cooldown timer. With recommendations of 12, 12, 5 and 4 inside the window and 12 copies running, the maximum is still 12, so no shrink happens. Only once the older, higher recommendations age out of the window does the maximum fall to 5 and the fleet shrink. A simple cooldown timer - "do not shrink for N minutes after any change" - is weaker, because it says nothing about what the loop recommended during those minutes. Note the asymmetry: the window applies on the way **down** only. Growth is usually allowed immediately, because being late to grow hurts users while being late to shrink only costs money. ## What a settling window cannot fix Damping treats oscillation caused by the loop's own feedback. It does not fix: - **A badly chosen signal.** If the measurement does not track the real constraint, a longer window just produces a smoother wrong answer. - **Genuinely bursty load.** If demand really does swing every few minutes, the count following it is correct behaviour, not flapping - and the right response is a higher floor or a buffer, not more damping. - **A skewed fleet.** If routing is uneven, the per-copy average can swing because work moved between copies, not because total load changed. Telling those apart is the diagnostic step: plot the *total* offered load beside the copy count. If total load is flat while the count oscillates, the loop is arguing with itself and damping is the answer. If total load is oscillating too, the loop is tracking reality and the question becomes how much standing capacity you want to hold - which is capacity planning, not a tuning problem. ## Platform variation Platforms differ in what they give you here: some expose separate windows for growth and shrink, some only a single cooldown, some let you declare a maximum rate of change per interval. The concept is constant - deliberately make shrinking slower and more conservative than growing - and the knob names and defaults are not, so check what yours actually implements rather than assuming a shape.

  • Why is the settling window applied on the way down and not on the way up?
    Because the two mistakes cost different things. Being late to add copies shows up as latency and errors for users; being late to remove them shows up on an invoice. So growth is usually allowed on the first qualifying reading, while shrinking waits for sustained evidence that the load really has gone.
  • How do you tell self-induced flapping from a genuinely bursty workload?
    Plot total offered load - requests per second or total queue depth - beside the copy count. If total load is flat while the count swings, the loop is reacting to its own actions and damping is the fix. If total load swings too, the count is tracking reality, and the question becomes how much standing capacity to hold rather than how to tune the loop.
  • Why can a flapping fleet serve worse than a fixed count sized for the peak?
    Because every cycle throws away warm capacity and then pays the full cold chain to get it back. Removal is immediate, addition takes placement, start-up, readiness and warm-up, so the fleet spends a large share of each cycle short of capacity or mid-start - while a static count would simply have been there.

A shower with a slow pipe. Turn the tap because the water is cold, and by the time it arrives you have overcorrected; the cure is to wait a few seconds before touching the tap again, not to turn it harder.

saying these in an interview costs you the question

  • Blames the workload for being bursty without checking whether total load moved
  • Treats a settling window as a timer rather than a high-water mark
  • Applies the same damping to growing as to shrinking
  • Assumes flapping only wastes money and never costs availability
  • Responds to a slow scale-up by making the loop more aggressive everywhere
  • Thinks a longer averaging window fixes a badly chosen signal