skip to content

questions

20

In an inline login-risk scorer, what is a per-device velocity counter and why is it kept continuously updated?

level: juniorimportance: must knowfreq 58%

answer

  1. a count, not a model
  2. keyed by one entity
  3. recent window, not all time
  4. written by the stream, read by the scorer
  5. one flat lookup beats a scan

basics

~20 s

A per-device velocity counter is a running count of that device's recent login attempts, held in a low-latency key-value store and folded forward by the login event stream, so the inline scorer reads it in one lookup instead of scanning history.

solid answer

~40 s

A velocity counter is a count of how often something happened to one entity inside a recent window — for a device, how many login attempts or how many distinct accounts it touched in the last hour. It is kept as a running value in a low-latency key-value store, updated as login events arrive on the event stream, and read by key when a credential submit is being scored. It is maintained rather than computed because the inline scoring budget is a few tens of milliseconds: aggregating a device's history at request time means a range scan whose cost grows with that device's activity, which is exactly the traffic an attacker generates. One `lookup(key)` is flat-cost no matter how busy the entity is.

go deeper

for a junior

Be able to say what the counter counts, what it is keyed by, and that it is read in one lookup because the scorer runs inside the login request.

for a middle

Explain the write path and the read path separately: events fold into the store continuously, the scorer only reads, and the time-to-live is what bounds storage.

for a senior

Show that you pick entity keys for the attack you expect, and that you carry a short and a long window per entity because the useful signal is the ratio between them.

for a principal

Frame the counter set as the cheapest risk surface the organisation owns, and the one that must keep working when the model does not.

## What a velocity counter is A **velocity counter** is the simplest useful feature on a risk-scoring path: the number of times something happened to **one entity** inside a **recent time window**. It is arithmetic, not modelling — nothing about it is learned, and the same counter can feed a model input, a rule, or an operations dashboard without changing meaning. On a consumer login path three counters carry most of the signal, and saying **which counter** you mean is load-bearing because they catch different attacks: | counter key | example feature | what it catches | |---|---|---| | per account | failed logins for this account in the last 5 minutes | someone guessing one victim's password | | per device | distinct accounts this device attempted in the last hour | one client replaying a stolen credential list | | per source network | attempts from this network address in the last minute | a single origin fanning out across many accounts | A bare "the velocity counter" is ambiguous in a design round, and the interviewer is usually listening for whether you name the entity key. ## Why it is maintained instead of computed The scorer runs **inside** the credential submit, before the session is granted, on a hard budget of a few tens of milliseconds. That budget is what rules out computing the count from raw events when the request arrives. | approach | cost at request time | behaviour under attack | |---|---|---| | scan this entity's event history | grows with how active the entity is | worst exactly when the entity is hostile | | read a maintained counter by key | one lookup, flat | unchanged by attack volume | The second row is the whole argument. An attacker is, by definition, the busiest entity in the system; a design whose read cost scales with activity gets slowest precisely on the requests that matter most. ## How it is kept up to date 1. Login and session events — attempt, success, failure, password reset, device first-seen — are published to an **event stream** as they happen. 2. A **stream processor** (or, for the shortest windows, the scoring service itself) folds each event into a counter in a **low-latency key-value store**, keyed by entity and window. 3. Each counter carries a **time-to-live** matched to its window, so an entity that goes quiet ages out of the store instead of being swept by a separate job. 4. The scorer reads the counters it needs by key, in parallel, inside the feature-fetch stage of its budget. ## What the window means - A counter with a TTL is an approximation of a **sliding window**: it forgets the whole window at once rather than attempt by attempt. - An exact sliding count needs the individual timestamps kept per entity, which costs more storage and a more expensive read — real systems pick per feature, and some keep both a short exact window and a long approximate one. - Long windows (24 hours, 30 days) describe an entity's **baseline**; short windows (60 seconds, 5 minutes) describe what it is doing **right now**. Most login scorers carry both, because the interesting signal is usually the ratio between them. ## What a velocity counter is not - **Not a decision.** A high count is an input; whether a login is allowed, challenged or declined is decided further down the path. - **Not a rule.** "More than 20 attempts in a minute" is a rule expressed over a counter, but the counter itself asserts nothing. - **Not unique to fraud.** The same maintained-count pattern shows up wherever a read must be cheap and the underlying event volume is unbounded; here it happens to be the cheapest risk signal there is. ## What it costs - **Read cost:** one lookup per entity key per scored request. Three keys means a fan-out of three, and the stage's latency is the **slowest** of them, not the sum. - **Storage:** roughly one entry per active entity per window, bounded by the TTL rather than by total history. - **Write cost:** one increment per event, off the request path when the stream processor owns it. - **A lag:** the value the scorer reads is as fresh as the pipeline that wrote it, which matters once an attacker moves faster than that pipeline. Because it is this cheap and this legible, a velocity counter is usually the first feature a risk system ships and the last one it would give up — a fraud team that loses its model can still run on counters, which is why they double as the degraded path when scoring fails.

  • Why do login scorers usually keep both a 60-second and a 30-day counter for the same entity?
    The short window says what the entity is doing right now; the long window says what is normal for it. Ten attempts in a minute means something very different from a device that averages two logins a day than from a shared family tablet. Most of the signal is the ratio, so the two windows are separate features rather than one.
  • What happens to a counter for an entity that stops appearing entirely?
    It expires with its time-to-live and the key disappears, so the store holds roughly the active entity set rather than all history. The scorer must therefore treat a missing key as a genuine zero-or-unknown case with an explicit default, not as an error, because first-seen devices are common on a consumer login path.

A doorman who keeps a tally sheet per visitor rather than re-reading the whole guest book each time someone knocks: the tally is instant no matter how often that visitor has knocked today.

saying these in an interview costs you the question

  • Treating a velocity counter as a decision rather than an input
  • Saying the count is computed by querying the login log when the request arrives
  • Leaving the entity key unnamed, so account, device and network blur together
  • Assuming one long window is enough, with no short-window counter
  • Expecting a missing key to be an error rather than a first-seen entity
open as a page

In a marketplace's fake-review defence, why send a middle band of risk scores to human reviewers instead of one allow/remove cutoff?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A single cutoff forces every borderline listing into an automatic allow or an automatic removal. The middle band is where the score is genuinely uncertain, so routing it to an analyst buys a cheap decision instead of a costly wrong one.

open as a page

Why does a card-not-present fraud model retrained on the last 30 days of checkouts learn that recent traffic is almost fraud-free?

level: middleimportance: must knowfreq 72%

basics

~10 s

A settled dispute confirms fraud weeks after the checkout, so recent rows carry no dispute yet. Joining absence-of-dispute to a legitimate training label marks immature fraud as good and deflates the recent fraud rate.

open as a page

In a refund-abuse decision layer, what changes if the rules run as a cascade before the model instead of entering it as features?

level: middleimportance: must knowfreq 74%

basics

~20 s

A cascade makes a matching rule an absolute veto and hides the traffic it decided from the model. Rule hits as features make the rule one weighable piece of evidence the model can discount. One is policy, the other is signal.

open as a page

When an allowlist entry and a hard-block rule both match one refund request, what decides the action the decision layer emits?

level: middleimportance: must knowfreq 66%

basics

~20 s

A precedence order declared in advance - not source order, not whichever rule was authored last. The layer collects the matches, resolves them by that documented ranking, emits exactly one action, and records which branch won.

open as a page

When the risk scorer misses its deadline on a login submit, should the login path fail open or fail closed?

level: middleimportance: must knowfreq 72%

basics

~20 s

Neither as a blanket default. Degrade down a ladder — a cached prior score for that account, then request-local rules, then an exposure-based default — and reserve a hard decline for high-exposure actions only, because a scorer timeout hits every login at once.

open as a page

A trust-and-safety team can work 600 seller-abuse cases a day; how do you place the review band's lower bound?

level: middleimportance: must knowfreq 56%

basics

~20 s

Work backwards from capacity. Pick the daily case volume the shift can sustain, take that many of the day's flagged candidates highest score first, and read the score at the bottom of that set as the bound, recomputed daily.

open as a page

Your fraud model is retrained only on checkouts it allowed — what does declining a transaction do to the next model's training data?

level: seniorimportance: must knowfreq 64%

basics

~20 s

A declined checkout is never charged, so almost nothing can dispute it and no outcome arrives. The declined region stays permanently unlabelled, and each retrain learns the fraud rate only among traffic the previous policy allowed.

open as a page

Why does a promotion and refund abuse decision layer keep hand-written rules after a model score is added?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Rules cover what is already known, instantly and deterministically: a promo-code exploit seen this week can be blocked today, and policy that must never bend is written rather than learned. The model covers the unenumerated remainder.

open as a page

A login submit gives risk scoring a 120 ms p99 budget; how do you divide it across feature fetch, model call and decision emit?

level: middleimportance: should knowfreq 62%

basics

~20 s

Give transport 20 ms and the scorer 100 ms: about 45 ms for a parallel feature fetch, 25 ms for the model call, 10 ms for deciding and handing off the emit, and a 20 ms reserve so the fallback still fits inside the same budget.

open as a page

An analyst cleared a checkout as legitimate and a dispute for unauthorised use settles against you six weeks later — which becomes the training label?

level: seniorimportance: should knowfreq 50%

basics

~10 s

The settled unauthorised-use dispute becomes the training label; the analyst verdict was a provisional stand-in. Keep both on an append-only record, because the disagreement rate measures the review policy rather than the model.

open as a page

What must a rules-plus-model decision layer record so a support agent can explain a denied refund months later?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The emitted action and its reason codes, every rule that matched with each rule's version, the model version and score band, which branch resolved the decision, and the input values it actually used - all written at decision time.

open as a page

Before a new hard-block rule acts on live refund requests, how do you find out what it would have done?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Run it in shadow: evaluate it on live traffic and record every match and the action it would have emitted, while the decision in force still stands. Replay over stored records first for cheap volume estimates.

open as a page

A stream processor updates login velocity counters with about two seconds of lag; why can a credential-stuffing burst be scored as low risk?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Because the counter the scorer reads describes the world as it was two seconds ago and never includes the attempt being scored. An attacker firing fifty attempts a second stays roughly a hundred attempts ahead of the count for the entire burst.

open as a page

A marketplace's fake-review flag volume triples during a seasonal spike: hold the score cutoff fixed, or hold daily review volume fixed?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Hold the volume. A fixed cutoff triples the queue's intake against an unchanged shift, so the backlog and time to decision blow out. Fixed volume raises the effective cutoff instead, trading a little more taken loss for decisions that stay on time.

open as a page

Your fraud team proposes allowing a random 2% of would-be-declined checkouts to earn real labels — how do you size and sign off that holdout?

level: principalimportance: should knowfreq 40%

basics

~20 s

Size it as an accepted-loss budget: sampled volume times the fraud share of declines times average loss per fraud. It buys an unbiased false-decline rate and training rows above the cutoff, so a risk or finance owner signs the budget, not an engineer.

open as a page

What must a marketplace's signed-off policy on wrongly removing honest sellers versus leaving fraudulent listings up actually state?

level: principalimportance: should knowfreq 38%

basics

~20 s

An accepted exchange rate between the two errors, owned by the business rather than the modelling team; a wrongful-removal ceiling stated as a measurable rate; whether the rate differs by seller tenure; the appeal path and its target; and who reviews it as the mix changes.

open as a page

Against an attacker who probes your checkout defences weekly, which of a fraud model's features lose their power first?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

The ones cheapest for the attacker to change — email patterns, device and browser details, shipping text, timing. Features tied to real-world cost, such as a money-out path or a physical delivery, age far more slowly.

open as a page

A mobile carrier's gateway fronts thousands of logins a minute from one network address; what does that do to your per-network velocity counter?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

It breaks the counter twice: the value is permanently high for thousands of ordinary subscribers, so it stops separating attack from normal traffic, and the single hot key concentrates reads and writes on one partition whose rising tail eats the feature-fetch stage for every login.

open as a page

In a seller-abuse review queue, what goes wrong when cases are worked in expected-loss order rather than arrival order?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

Low-value cases starve. Each day's high-loss arrivals jump the queue, so a cheap case can sit for weeks, breach its time-to-decision target and eventually expire unworked while an honest seller waits in limbo.

open as a page