A mobile carrier's gateway fronts thousands of logins a minute from one network address; what does that do to your per-network velocity counter?
answer
- shared egress, many real users
- the count measures popularity
- compare to the network's own baseline
- one key, one partition
- spread writes, cap the read fan-out
basics
~20 sIt breaks the counter twice: the value is permanently high for thousands of ordinary subscribers, so it stops separating attack from normal traffic, and the single hot key concentrates reads and writes on one partition whose rising tail eats the feature-fetch stage for every login.
solid answer
~50 sThe first failure is about the signal. A shared egress address means the per-network counter is large for every legitimate user behind it, so a raw count keyed on that address says "attack" for an entire carrier's subscriber base. The fix is to stop reading the raw number: compare the count against that network's own established baseline, or use a shape like distinct accounts per attempt, and demote the network feature when the address is known to be shared. The second failure is about serving. One key taking thousands of operations a second concentrates on a single partition of the low-latency key-value store, and that partition's tail latency is paid by **all** scored logins, not just the ones behind the gateway. Shard the counter across sub-keys summed on read, or aggregate in process and flush periodically, accepting bounded staleness on that one feature.
go deeper
Know that many unrelated people can share one source network address, so a count keyed by that address is not a count of one person's behaviour.
Explain both consequences: the feature stops discriminating for that address, and one key concentrated on one partition becomes a store hotspot.
Give concrete fixes for each — normalise against the network's own baseline or use a distinct-accounts shape, and spread writes across sub-keys or flush in-process aggregates — and say which you would do first.
Frame it as a fairness and availability question as much as a risk one: a feature that penalises whole populations behind shared infrastructure is a product problem the risk owner has to answer for.
## One entity key, two unrelated failures A shared egress point — a carrier gateway, a corporate proxy, a large campus network — breaks a per-network velocity counter in two ways at once, and they need different fixes. Candidates who see only one usually see the latency problem and miss the signal problem, which is the more expensive of the two. ## The signal failure A velocity counter keyed by source network is evidence about *whoever is behind that address*. When the address fronts one client, that is a useful proxy for one actor. When it fronts a hundred thousand subscribers, the count measures the carrier's popularity. What goes wrong follows directly: - **The feature stops discriminating.** Its value is high for hostile and benign traffic alike from that address, so it carries almost no information for those requests. - **Anything keyed off the raw count over-triggers.** A challenge condition written against attempts-per-network puts an entire carrier's users through step-up, which reads to the business as an outage with a different name. - **It is worst for the users least able to avoid it.** Mobile networks and large shared workplaces concentrate exactly the customers a consumer service cannot afford to treat as suspicious. Fixes are about **normalisation and shape**, not about a bigger number: - compare the current count to that network's **own established volume** rather than to a global constant, so the signal is "unusual for this network", not "large"; - prefer shapes that survive sharing — **distinct accounts per attempt**, the failure-to-success ratio, or the fraction of attempts using credentials never seen from that network before; - maintain a set of **known shared egress ranges** and demote or drop the network feature for them, letting the per-device and per-account counters carry the decision; - remember that any allowlist of shared ranges is itself an attack surface, since sitting behind a demoted network is now valuable. ## The serving failure The same address is also a single key in the counter store, and a single key lives on a single partition. Thousands of operations a second against one key produce a genuine hotspot: - writes to that key **contend**, so update latency rises for it specifically; - the partition serving it also serves other keys, so **unrelated lookups slow down too**; - the feature-fetch stage waits on its slowest branch, so a hot partition's p99 is paid by **every** scored login, whether or not it is behind that gateway; - and the inline budget has no slack for it, so the visible symptom is a rise in degraded decisions across the whole login path. The fixes trade exactness for spread: 1. **Shard the key.** Write to `networkKey:0 … networkKey:N-1` picked at random and sum the N values on read. Writes spread across partitions; reads cost N lookups instead of one, so N stays small and is applied only to keys identified as hot. 2. **Aggregate in process, flush periodically.** Each scorer instance keeps a local delta for the hot key and flushes it every few hundred milliseconds. Store traffic drops by the number of attempts per flush, at the cost of the counter lagging by one flush interval. 3. **Read-cache the hot key briefly.** For a key taking thousands of attempts a minute, one more attempt barely moves the value, so a very short time-to-live on a cached read is nearly free in accuracy. 4. **Detect hotness rather than hard-coding it.** Track per-key operation rates and promote a key into the sharded or cached treatment automatically, because which networks are hot changes with the traffic. ## Why the two failures get confused They arrive together and look like one incident: latency rises, degraded decisions rise, and challenge volume rises. But treating it purely as a capacity problem leaves a feature in production that recommends challenging a whole carrier, and treating it purely as a signal problem leaves a hot partition that will re-appear with the next large shared network. In a design round, saying which of the two you would fix first — and that the answer is the signal, because it is silently wrong even when the store is fast — is usually the point of the question.
- Why does the hot key slow down logins that have nothing to do with that network?The key lives on one partition of the store, and that partition also serves unrelated keys. Once it is saturated its tail latency rises for everything it holds, and since the feature-fetch stage waits on its slowest parallel branch, any request unlucky enough to touch that partition pays it. The symptom is a broad rise in degraded decisions, not a localised one.
- What does sharding one counter across N sub-keys cost?Reads become N lookups that must be summed, which multiplies the read cost of a stage that is already the largest slice of the inline budget. That keeps N small and makes it worth applying only to keys measured as hot, ideally promoted automatically rather than configured by hand.
- Is maintaining a list of known shared networks a complete answer to the signal problem?No. The list is always incomplete and always stale, and being on it becomes valuable to an attacker, since it demotes a feature. It works as one input alongside normalising against the network's own baseline and using shapes such as distinct accounts per attempt, which degrade gracefully when the list is wrong.
saying these in an interview costs you the question
- Reading a raw per-network count as evidence about one actor
- Challenging every user behind a shared gateway and calling it risk control
- Seeing only the hotspot and leaving the feature meaningless
- Assuming a hot key slows down only requests that read it
- Sharding every counter rather than the few measured as hot
- Trusting a shared-network allowlist as a complete defence