skip to content

Launch morning traffic is 30x normal: what does the request-rate graph alone prove about who is sending it?

level: juniorimportance: must knowfreq 70%

answer

  1. a count, not an identity
  2. same curve, two very different senders
  3. the announced hour is the disguise
  4. diversity at the edge, shape at the app
  5. both wrong calls cost money today

basics

~20 s

Only that requests arrived. A rate count carries no sender identity, no intent and no outcome. A flood shaped to look like growth and a real launch draw the same curve, which is exactly why an attacker picks that hour.

solid answer

~50 s

It proves arrivals arrived at the measurement point, and nothing else. Request rate is a scalar: there is no sender identity in it, no intent, and no evidence that anything was served usefully. A successful launch and a flood timed to hide inside a launch produce the same shape, and that cover is the whole reason an adversary picks the window. To separate them you need two other measurements. At the edge: how the sources spread across networks and how many distinct client fingerprints those sources present. At the application: whether the arrivals behave like sessions at all — do they come back for a second page, carry state forward, ever complete anything. Acting on the rate alone is a coin flip with expensive sides: arm mitigation and you turn away real buyers on the one day they came; hold and you serve a flood at full price.

go deeper

for a junior

Be ready to say plainly that a rate graph counts arrivals and carries no sender, no intent and no outcome, and to name at least one other measurement you would take before acting.

for a middle

Explain why baseline-relative anomaly detection is weakest on an announced launch day, and describe concretely how source diversity and session shape are computed from logs you already hold.

for a senior

Show that you price both errors before you act: what arming costs in lost genuine buyers, what holding costs in served capacity and bill, and which of those is recoverable.

for a principal

Own the framing that a capacity signal is not an identity verdict, and that on-call teams need a pre-agreed evidence list for launch days rather than a threshold on a graph that everyone knows is uninformative.

## What a rate graph contains, and what it does not A requests-per-second or bits-per-second line is an aggregate count over a time window. Everything that made each request distinguishable — the source, the client software, the path asked for, whether anything downstream succeeded — has been summed away before the point is plotted. So the honest reading of a 30x spike is: *the measurement point counted thirty times as many arrivals as it usually does in this window*. That is a real, useful fact. It is not a claim about who sent them, why, or whether serving them is worth anything. This matters here because both stories that explain the spike are ordinary: - A launch went well. Marketing pushed at a fixed hour, a link travelled, and tens of thousands of genuine people arrived within a few minutes. Real launch curves are not gentle; they can be near-vertical. - Someone aimed a flood at the same hour. The adversary's advantage is that your own expected curve is the disguise. Anomaly detection that works by comparing today's rate against a historical baseline is at its weakest precisely when you have publicly announced that today will not look like history. ### Why 'it went up too fast to be organic' is the wrong answer The common junior answer is that organic growth is gradual, so a vertical edge implies an attack. It does not. A promoted launch, a link on a large aggregator, a push notification to an installed app base, or a broadcast mention all produce step functions. Conversely, a competent flood can be ramped slowly on purpose. Rate slope is a hint about *scheduling*, not about *who*. ### What the rate does buy you Rate is still the right thing to alarm on for capacity: it tells you which box in the path is about to fill, and how long you have. That is a legitimate, separate use. The error is promoting a capacity signal to an identity verdict. ### The measurements that actually discriminate Two classes of evidence survive the aggregation and are computable while traffic is still arriving: 1. **Diversity at the edge.** Not just how many distinct source addresses, but how the sources spread across distinct networks *and* how many distinct client fingerprints those sources present. A real crowd is a messy population of browser versions, mobile and desktop, different negotiated protocol versions. A single tool run from many places is broad in addresses and narrow in fingerprints. 2. **Shape at the application.** Do the arrivals behave like sessions? Real visitors fetch a second thing, return state they were given, and some fraction of them finish something. Traffic sent to consume capacity tends only ever to arrive. Neither of these is a rate. Both can be produced quickly from logs you already keep. ### The price of getting it wrong, in both directions The reason this is an interview question rather than a lecture is that both errors cost real money on the same morning: | Call | If you are wrong | | --- | --- | | Arm mitigation | Genuine buyers meet a challenge, a queue, or a drop on the one day they were going to buy. The revenue does not come back later, and marketing is on the bridge watching. | | Hold | You serve a flood at full price: origin capacity, egress bandwidth, per-request compute, and whatever autoscaling bill that generates — and you may still fall over, later, with less time to react. | So the answer an interviewer wants is not merely 'the graph proves nothing'. It is: the graph proves arrivals; here are the two other measurements I would take before acting; and here is what each wrong call costs, stated out loud, before I make it. ### A related trap: your own clients A third story explains many spikes and is neither growth nor attack: your own client software retrying. A mobile app that fails a call and immediately re-issues it turns one degraded backend into a self-inflicted arrival storm. On a rate graph that is indistinguishable from the other two, which is one more reason the graph is a starting point rather than a verdict.

  • If the rate graph proves so little, why keep alarming on it at all?
    Because it is the right signal for a different question: capacity. Rate tells you which component in the path is filling and roughly how long you have before it does — which is what decides whether you have ten minutes to gather evidence or ten seconds. The mistake is promoting a capacity signal into a verdict about who is sending the traffic.
  • Your own mobile client is retrying hard against a slow backend. How does that show up next to the other two stories?
    As a spike that is indistinguishable on rate alone, which is the point. It usually separates on the other axes: the sources are your real installed base, so fingerprints look like your app rather than one tool; the arrivals often carry valid prior state; and the volume tracks your own error rate rather than an external event. Treating it as an attack and blocking it locks out your own users.
  • Marketing asks on the bridge whether the spike means the campaign worked. What do you say?
    That arrivals are up 30x and that arrivals are not the same as buyers. The number that answers their question is completions — purchases, sign-ups, whatever the launch was for — per minute, and that same number is one of the measurements that will tell us whether the traffic is real. Promise them that figure rather than the rate.

A turnstile counter tells you three thousand people came through the gate. It cannot tell you whether they are fans or a crowd sent to jam the entrance, and the ticket office is about to make an expensive decision on that number alone.

saying these in an interview costs you the question

  • Treats a vertical spike as proof of an attack
  • Says organic growth is always gradual
  • Reads bits per second as evidence of who sent the traffic
  • Arms a blocking mitigation on the rate graph alone
  • Ignores that turning away real buyers has a price too

context