skip to content

During a suspected flood, why is 'thousands of source addresses across hundreds of networks' not evidence of an attack?

level: middleimportance: should knowfreq 55%

answer

  1. both stories predict the same spread
  2. NAT breaks the reverse inference too
  3. reach is not intent
  4. count clients, not addresses
  5. one fingerprint, six hundred networks

basics

~20 s

Because a genuine crowd is distributed too. Address spread measures reach, not intent. The discriminating structure is the joint one: many sources presenting very few distinct client fingerprints means one client run from many places, not many people.

solid answer

~50 s

Address diversity on its own is symmetric — a real launch crowd also arrives from thousands of addresses across hundreds of networks, and carrier-grade NAT and mobile networks collapse thousands of real users behind a handful of addresses, so low diversity is not evidence of a launch either. What separates them is the pairing. Alongside source spread, count distinct client fingerprints: the shape of the TLS ClientHello, the negotiated protocol and ALPN, header ordering, the declared user agent. A real population is messy — dozens of browser builds, mobile and desktop, several protocol versions. A flood is usually broad in addresses and narrow in fingerprints, because one tool is being run from many places. The price of this measurement is that you must own a vantage that terminates or inspects the handshake, and a careful adversary can randomise the fingerprint to blur it — so treat narrow fingerprint diversity as strong evidence and wide diversity as no evidence either way.

code

text · 6 lines
text
src=203.0.113.44   asn=AS64500  GET /launch  200  sni=shop.example  alpn=h2  tls_fp=e7d1..9c  ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
src=198.51.100.9   asn=AS64501  GET /launch  200  sni=shop.example  alpn=h2  tls_fp=e7d1..9c  ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
src=192.0.2.130    asn=AS64502  GET /launch  200  sni=shop.example  alpn=h2  tls_fp=e7d1..9c  ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
src=203.0.113.201  asn=AS64500  GET /launch  200  sni=shop.example  alpn=h2  tls_fp=e7d1..9c  ua="Mozilla/5.0 (Windows NT 10.0; Win64; x64)"
...
41,000 arrivals in 60s | 8,900 distinct src | 640 distinct asn | 1 distinct tls_fp | 1 distinct ua | 0 requests for any second path

go deeper

for a junior

Know that a real crowd is also spread across many networks, so counting source addresses cannot by itself tell an attack from a successful launch.

for a middle

Be able to explain what a client fingerprint is made of — handshake structure, negotiated protocol, header order, declared agent — and why broad sources with narrow fingerprints means one client run from many places.

for a senior

Show you know the inference is one-directional and that the collateral of fingerprint-keyed action is a whole client population, so you name whose fingerprint it is before you act on it.

for a principal

Own the fact that this evidence only exists if the handshake fields were being logged and retained, which is a storage and vantage decision taken long before launch day.

## Why address spread is a symmetric signal The instinct is understandable: distributed denial of service is distributed, so a long list of distinct sources feels like proof. It is not, because the *other* explanation is equally distributed. A promoted launch draws from every consumer network in the market — thousands of addresses, hundreds of autonomous systems, arriving in the same minutes. Both stories predict the same observation, so the observation does not discriminate. The reverse trap is worse. Low diversity is not evidence of legitimacy either: - Carrier-grade NAT puts tens of thousands of real mobile users behind a small pool of addresses. - A corporate estate egresses through a handful of proxies. - Traffic that reaches you via an intermediary — a caching layer, a partner, a link-preview fetcher — arrives as that intermediary, and unless the real client address is carried in a header you trust for a specific reason, you cannot see behind it. So the count of distinct addresses tells you about *reach*, and reach is not intent. ## The pairing that does discriminate The useful measurement is joint: source spread against **client fingerprint cardinality**. A client fingerprint is what the client reveals about itself before any application data is exchanged and in the request it then sends — the ordering of cipher suites and extensions in the TLS ClientHello, the supported groups, the ALPN it offers, the HTTP version it negotiates, the order and casing of its headers, and the user agent it declares. None of that is secret and none of it is authenticated; it is simply structure, and different client software produces different structure. A real crowd is a population sample. It carries dozens or hundreds of distinct fingerprints because it is dozens of browser builds on several operating systems, with a large mobile fraction, some of them proxied, some of them old. A flood is often one program compiled once and run from many places. The addresses fan out; the fingerprint does not. **Thousands of sources presenting one fingerprint is one client, distributed** — and that is a fact about the traffic that no amount of source spread can explain away. ## Where the pairing breaks, and you must say so An interviewer will push here, and the honest answer has three caveats: 1. **A single legitimate fingerprint across many addresses is normal in one important case:** your own mobile application. If the launch included a push notification to an installed base, the resulting arrivals genuinely are one client run from a hundred thousand places. That cohort should look identical to a botnet on this measure and must be excluded by knowing your own client, not by the metric. 2. **The adversary can blur it.** Randomising extension order, rotating user agents, or driving real browser engines raises fingerprint cardinality on purpose. Encrypted Client Hello and other handshake privacy work also reduce what a passive observer sees. So the inference is one-directional: narrow fingerprint diversity across broad source diversity is strong evidence; wide diversity is not evidence of legitimacy, only absence of the cheap tell. 3. **You must own the vantage.** Fingerprint cardinality is only computable where the handshake is visible — the terminating edge, or an inspection point that records ClientHello structure. Flow records carry the five-tuple, byte and packet counts and timestamps and no payload at all, so they can show that bytes moved and never which client moved them. If your logging pipeline drops those fields to save storage, this evidence does not exist when you need it, and that storage decision was made months ago by someone who was not on this bridge. ## The comparison problem Diversity is most useful compared against your own history: how many fingerprints per thousand sources do you normally see on this path? In a merged estate, that history may not exist in a comparable form — two front doors, two logging stacks, two field sets, no shared denominator. Structural facts survive that better than ratios do. 'One fingerprint across 640 networks' needs no baseline to be alarming; '40x normal' needs a baseline you may not have. ## The price Calling it wrong in the direction of blocking has a specific shape here: fingerprint-based mitigation collateral is not random. If you challenge or drop a fingerprint, you drop *everyone using that client build* — which can mean every user on one mobile OS version, or every user behind one enterprise proxy, or your own app. That is a defensible action only when you can name whose fingerprint it is and accept losing them, and it is why arming reversibly, with a bypass for sessions that have already proved themselves, matters more than picking a clever discriminator.

  • Your own mobile app got a launch push and now shows one fingerprint across a hundred thousand addresses. How do you avoid blocking it?
    By knowing your own client's fingerprint before launch day and holding it on an explicit allow list, rather than hoping the metric will spare it. That is a pre-launch task: record what your app and your official web client look like on the handshake, and treat any mitigation keyed on fingerprints as needing that list to exist first.
  • What can you conclude about client software from flow records alone?
    Essentially nothing. A flow or IPFIX record carries the five-tuple, byte and packet counts and timestamps, and no payload at all, so it can show that bytes moved between two endpoints and never which client moved them. Fingerprint evidence requires a vantage that sees the handshake, which is a logging and placement decision made long before the incident.
  • An adversary randomises user agents and extension ordering. Does the measurement become useless?
    It loses its cheap form, not all of it. Wide fingerprint diversity stops being evidence of legitimacy — it just means the tell was hidden — so you fall back to shape at the application: whether the arrivals ever fetch a second thing, carry state forward or complete anything. Randomising a handshake is cheap; simulating a real session's forward path is not.

Ten thousand letters postmarked from every town in the country tells you nothing; ten thousand letters in identical handwriting, postmarked from every town in the country, tells you a great deal.

saying these in an interview costs you the question

  • Treats many distinct source addresses as proof of a botnet
  • Treats few source addresses as proof of legitimate traffic
  • Forgets that carrier-grade NAT hides thousands of real users
  • Claims flow records reveal which client software connected
  • Blocks on a fingerprint without knowing whose client it is

context