skip to content

Your two edge stacks disagree on a launch-day spike — 40x normal versus 3x — with ten minutes to decide. How do you call it?

level: seniorimportance: must knowfreq 45%

answer

  1. two ratios, two denominators
  2. no shared history to settle it
  3. ask what shape, not how much
  4. choose the recoverable wrong answer
  5. state the falsifier before arming

basics

~20 s

Stop trying to reconcile the multiples; they are ratios over different denominators and no shared history exists to settle them. Switch from how much to what shape, act reversibly on the launch path only, and say out loud what would make you disarm.

solid answer

~50 s

The two numbers are not in conflict about reality, they are ratios over different denominators — different populations behind each front door, different definitions of a request, possibly sampled differently, and after an acquisition no shared history to normalise them against. Reconciling them takes longer than the launch clock allows, so I stop asking how much and ask what shape: source spread against client fingerprint cardinality on each stack, and whether arrivals are turning into sessions with a forward path. Those are structural, computable in minutes, and need no common baseline. Then I pick an action whose wrong side is recoverable: a challenge or graduated admission rather than a hard drop, scoped to the launch path only, with an explicit bypass for clients that already hold valid state, and one person watching completions per minute rather than requests per second. Before arming I state the falsifier on the bridge and the revert trigger, and I write down the call, the evidence and the time.

go deeper

for a junior

Understand that a multiplier like 40x is a ratio against a baseline, so two systems can report very different multiples for the same traffic.

for a middle

Be able to list the concrete reasons two stacks disagree — different populations, different definitions of a request, sampling, window alignment — and why structural evidence needs no shared baseline.

for a senior

Show the operating discipline: change the question to shape, pick the reversible action scoped to the launch path, bypass clients that already hold valid state, and state the falsifier and revert trigger before arming.

for a principal

Own the framing that the two telemetry planes are an unpaid integration debt, and that the launch-day decision quality is set by choices about request definitions, fingerprint records and retention made months earlier.

## Why the two numbers disagree, and why that is not a puzzle to solve now After a merger you have two of everything: two front doors, two logging stacks, two sets of field definitions, and one joint launch pointed at both. A multiplier like *40x normal* is a ratio, and a ratio has a denominator. The two stacks can differ in every part of it: - **Population.** The two properties serve different customer bases with different daily rhythms. A quiet baseline makes any spike look enormous. - **What counts as one request.** One stack may count application requests behind a reused connection while the other counts connections, or object hits including every subresource. That alone is easily an order of magnitude. - **Sampling.** If one plane samples flows or requests and scales up, its short-window ratio is noisier and can be badly off at a step change. - **Window and alignment.** A one-minute rate against a five-minute rolling mean is a different number from a five-minute rate against a same-hour-last-week mean. - **History.** The whole point of a merged estate is that there is no shared history to compare today against, so neither denominator is authoritative. Reconciling all that is a week of work. You have ten minutes and a clock nobody controls. ## Change the question The move a senior engineer makes is to stop asking *how much* and start asking *of what shape*, because shape needs no shared baseline: - On each stack independently: how do sources spread across networks, and how many distinct client fingerprints do they present? Broad sources with one fingerprint is one client run from many places, and that statement is true without any history. - At the application: are arrivals turning into sessions with a forward path — a second request, state returned, anything completed? - Do the two stacks agree on *shape* even though they disagree on multiple? If the 3x stack shows the same one-fingerprint structure as the 40x stack, the difference between them is measurement, not phenomenon. If only one shows it, you have located the target and can scope the action. Structural facts are portable across mismatched telemetry. Ratios are not. ## Act reversibly, and price both sides out loud Both errors are expensive and in opposite currencies: turn away genuine buyers on the single day they arrived, or serve a flood at full price in origin capacity, egress and bill while risking the outage anyway. Since you cannot eliminate the risk, choose the action whose wrong side you can undo: 1. **Graduated rather than binary.** A challenge or an admission gate degrades a wrong call into friction; a hard drop turns it into lost customers you never see again. 2. **Scoped.** Apply to the launch path, not the whole front door, so the rest of the business is untouched by your guess. 3. **Bypassed for the already-proved.** Clients holding valid state you issued, and authenticated sessions, skip the gate. This is what protects the cohort you most cannot afford to lose. 4. **Known-client allowance.** Your own mobile application legitimately presents one fingerprint from a hundred thousand addresses; if you did not record what it looks like before today, you cannot safely act on fingerprints at all. 5. **A revert trigger with a name against it.** Not *we will keep an eye on it* but a stated condition and a person who will act on it. ## Say the falsifier before you act The habit that separates a good launch-day owner from a lucky one is announcing, before arming, what evidence would change the call in the next ten minutes. For example: *if fingerprint cardinality rises toward the population we normally see and completions start tracking arrivals, I disarm within five minutes.* Marketing is listening, and this converts an argument about instinct into a small number of checkable claims. It also protects you afterwards, because the decision will be re-litigated whichever way it went. ## Watch the outcome, not the arrivals Once armed, the number that tells you whether you were right is not requests per second — it is whether people are finishing what the launch was for, per minute. Arrivals are what the adversary controls; completions are what you were defending. If completions rise after arming, you removed contention. If they fall, you are blocking the crowd, and the revert trigger exists for exactly that. ## Afterwards Write the call, the evidence, and the timestamps down at the time, while the traffic is still arriving. That is not incident paperwork; it is the only record that will exist of what was actually visible in the ten minutes, and the reconciliation work the two stacks obviously need — one agreed request definition, one comparable denominator, a recorded fingerprint for your own clients — is the follow-up this morning has just paid for.

  • Both stacks show the same one-fingerprint structure but very different multiples. What does that tell you?
    That the disagreement is measurement, not phenomenon — the same traffic is hitting both doors and the denominators differ. It also means the action has to cover both front doors, because scoping to the stack with the alarming number would simply move the load to the other one.
  • Marketing asks you not to arm anything because the campaign cost a fortune. How do you answer on the bridge?
    By naming both prices in their currency: a challenge costs some friction on the launch path, a hard drop costs buyers, and doing nothing risks the page being down for everyone during the window they paid for. Then offer the reversible option with a stated revert trigger, so the decision is not permanent and they can hear exactly what would undo it.
  • You arm a challenge and arrivals stay flat while completions rise. What happened?
    You most likely removed contention — genuine users who were losing to the flood are now getting served. Flat arrivals with rising completions is the signature of a correct call. If instead completions had fallen with arrivals, you were gating the crowd and the revert trigger should fire.
  • Which follow-up work does this morning justify, regardless of which way the call went?
    One agreed definition of a request across both stacks with a comparable denominator, a recorded fingerprint for your own official clients so they can be allowed explicitly, and retention of the handshake and correlation fields the shape measurements depend on. All three are cheap now and impossible to arrange in ten minutes.

Two thermometers in different rooms disagree about how far above normal today is. You do not calibrate thermometers with the house on fire; you go and look at whether there is smoke.

saying these in an interview costs you the question

  • Spends the window trying to reconcile the two multipliers
  • Treats the larger multiple as the true one
  • Arms a hard drop across the whole front door
  • Acts without a stated revert trigger or owner
  • Judges the outcome on requests per second rather than completions

context