skip to content

A stream processor updates login velocity counters with about two seconds of lag; why can a credential-stuffing burst be scored as low risk?

level: seniorimportance: should knowfreq 47%

answer

  1. the read is a snapshot of the past
  2. this attempt is not in it
  3. rate times lag equals invisible attempts
  4. write the short window on the request
  5. attempt id keeps retries honest

basics

~20 s

Because the counter the scorer reads describes the world as it was two seconds ago and never includes the attempt being scored. An attacker firing fifty attempts a second stays roughly a hundred attempts ahead of the count for the entire burst.

solid answer

~50 s

Two separate gaps open at once. The **pipeline lag** means the value in the low-latency key-value store reflects events the stream processor has already folded in, so anything from the last couple of seconds is invisible. On top of that, the **current attempt is never in the count**, because it has not been published and processed yet — an off-by-one that matters most at the start of a burst, when the true count is one and the read count is zero. At fifty attempts a second, two seconds of lag hides about a hundred attempts, which is the whole attack. The fix is to stop relying on the stream for the short window: have the scoring path itself perform an idempotent increment in the counter store and read back the post-increment value, keeping the streamed aggregate only for long windows where seconds do not matter.

code

pseudocode · 14 lines
pseudocode
onLoginAttempt(attemptId, accountKey):
    // long window: folded in by the stream, seconds of lag are immaterial here
    slow = lookup(onlineStore, accountKey).attempts30d

    // short window: written and read by the scorer itself, so it sees this attempt
    fast = addAndCount(counterStore,
                       key    = accountKey + ":60s",
                       member = attemptId,        // same id twice = no change
                       ttl    = 60s)

    features.attempts30d = slow    // baseline for this account
    features.attempts60s = fast    // includes the attempt being scored
    // kept as two features; never summed, or the overlap is counted twice
    return features

go deeper

for a junior

Understand that a counter read during a request reflects events already processed, so very recent activity, including the current attempt, may not be in it yet.

for a middle

Separate the two gaps — pipeline lag and the uncounted current event — and be able to multiply attack rate by lag to say how much is invisible.

for a senior

Design the split: short windows written synchronously on the scoring path with idempotent increments, long windows left to the stream, and the two kept as distinct features.

for a principal

Weigh a write on the critical path against the loss it prevents, and decide how much inline budget and store capacity the organisation will spend to see a burst while it is happening.

## Two different gaps, not one When the scorer reads a velocity counter it gets a number that is wrong in two independent ways. 1. **Pipeline lag.** The event has to be published to the stream, picked up by the stream processor, folded into the keyed counter and written to the low-latency key-value store. That round trip is typically hundreds of milliseconds to seconds, and it is a distribution, not a constant — under load the lag grows exactly when attack volume is highest. 2. **The attempt being scored is not in the count.** Even with a zero-lag pipeline, the event for *this* credential submit is generated by the very request the scorer is answering. It cannot already be aggregated. This is an off-by-one on every single request, and it is most damaging at the first attempt of a burst, where the true count is one and the read count is zero. ## Why a burst outruns the pipeline Put numbers on it. An attacker replays a stolen credential list at **50 attempts per second** against one account list, and the pipeline lag is **2 seconds**: - attempts invisible to the scorer at any moment: about **50 x 2 = 100**; - if the rule the risk layer cares about fires at, say, a few dozen attempts in a minute, the burst can be substantially complete before the counter ever crosses it; - worse, the attacker controls the rate. A burst deliberately shorter than the lag is scored entirely against a counter that says zero. This is why "only two seconds" is not a reassuring number on a risk path. Seconds of lag are immaterial for a 24-hour counter and fatal for a 60-second one. ## The fix: write on the scoring path The short window must not come from the stream at all. The scorer performs the increment itself, synchronously, against the counter store, and reads back the value **after** its own increment: - the count now always includes the attempt being scored, removing the off-by-one; - there is no pipeline between the event and the count, removing the lag; - the cost is one extra write inside the feature-fetch stage of the inline budget — real, but a single-key write to a low-latency store, and it can be issued in parallel with the other lookups. ## Composing the two sources | window | written by | lag | why it is acceptable | |---|---|---|---| | 60 seconds | the scoring path, on the request | none | needed to see a burst in progress | | 24 hours / 30 days | the stream processor | seconds | seconds are noise against a day | The two are kept as **separate features**, not added together. Once the pipeline catches up, the long aggregate contains the same attempts the short counter already counted, so summing them double counts the overlap. Keeping them distinct is both simpler and more informative — the interesting signal is usually the short window read against the entity's long-window baseline. ## Idempotency An increment on the request path has to survive retries. A client retry, a load-balancer retry or an internal retry of the scoring call must not add to the count twice, or the system manufactures the burst it is trying to detect: - carry an **attempt identifier** generated once by the login service, not by the scorer; - make the counter a set of recent attempt identifiers with a time-to-live matched to the window, so the count is the set's cardinality and adding the same identifier twice is a no-op; - or write the identifier to a short-lived dedupe key and only increment when the write is the first. The set form costs more memory than a plain integer, which is why it is usually reserved for the short windows where exactness matters. ## What this does and does not solve - It **does** make the score reflect the burst in progress, including the current attempt. - It **does** keep the expensive long-window aggregation off the request path, where it belongs. - It **does not** help if the attacker spreads across entity keys faster than any single counter can see — one attempt per account across ten thousand accounts leaves every per-account counter at one. That is what the per-device and per-network counters exist for, and why a risk system carries several keys rather than a better version of one. - It **does not** remove the need to handle a counter-store failure: if the synchronous increment times out, the scorer is back on its deadline discipline and its degraded path, scoring with the feature missing rather than waiting.

  • Why not just shorten the pipeline lag instead of writing the counter on the request path?
    Lag can be reduced but not removed, and it is a distribution that widens under exactly the load an attack creates. More fundamentally, no pipeline can include the attempt currently being scored, because that event does not exist until the request is handled. Shortening the lag narrows the gap; writing on the path closes it.
  • What breaks if the synchronous increment is not idempotent?
    Any retry — client, proxy or internal — inflates the entity's count, so the system fabricates velocity that never happened and challenges or blocks users on its own duplicate traffic. Keying the increment on an attempt identifier minted once by the login service makes a repeated write a no-op and keeps the count equal to the number of real attempts.
  • The attacker sends one attempt per account across ten thousand accounts. Which counter catches it?
    Not the per-account one, which stays at one everywhere. The per-device and per-source-network counters do, because the breadth that defeats a per-account view concentrates on whatever the attempts share. This is the standard argument for carrying several entity keys rather than trying to make one counter smarter.

saying these in an interview costs you the question

  • Believing the streamed counter already includes the attempt being scored
  • Calling two seconds of lag negligible on a sixty-second window
  • Adding the short-window counter to the long aggregate and double counting
  • Incrementing on the request path with no attempt identifier
  • Assuming a faster pipeline would remove the gap entirely
  • Relying on one entity key to catch an attack spread across many