skip to content

A credential-stuffing run converts 0.2% of pairs into logins — why is that profitable, and what kills it?

level: seniorimportance: should knowfreq 45%

answer

  1. the operator's unit is not the percentage
  2. the input costs nothing
  3. proxies and solvers are line items
  4. cost per valid session
  5. shrink the overlap, not the request rate

basics

~20 s

Because the corpus is nearly free and each valid session resells. Two thousand hits from a million pairs clears the proxy and challenge-solving bill. Only shrinking how many of your accounts appear in that corpus changes the arithmetic.

solid answer

~50 s

The operator does not optimise hit rate; it optimises **cost per valid session**. The corpus of `address:password` pairs is aggregated from years of other organisations' dumps and is cheap or free. The real costs are request-side: rented residential source addresses so no per-source counter accumulates, and per-request fees for solving whatever challenge sits in front of the login. Against a public customer login, 0.2% of a million pairs is two thousand working sessions, and the resale value of even a low-tier consumer session comfortably exceeds a fraction of a cent per request. Everything you can do to the request path — per-source throttling, challenges, device checks — raises the price per attempt, which the operator has already budgeted and passes to a solver. The one term it cannot buy around is the hit rate itself, and the hit rate is the fraction of your accounts whose current password appears in the corpus. Screening every password against known-compromised values at set time is the change that moves it.

go deeper

for a junior

Know that a very low hit rate can still be profitable because the stolen pairs cost the attacker almost nothing and each working sign-in has resale value.

for a middle

Break the run's cost into its parts — corpus, rented source addresses, per-solve fees — and explain why the operator optimises cost per valid session instead of hit rate.

for a senior

Be ready to say which proposed changes tax the run and which remove it, and to redirect a team from request-path controls to measuring and shrinking the reuse overlap.

for a principal

Own the argument that an attack with a free input cannot be priced out, only starved of yield, and be able to justify the spend on screening in those terms to an owner who wants a request-path fix instead.

## Price the run, not the percentage A 0.2% hit rate sounds like failure to an engineer and reads as a healthy margin to the operator, because the two are measuring different things. The engineer is measuring success per attempt. The operator is measuring **cost per valid session**, and it drives that number down by attacking the denominator — the cost of an attempt — rather than the numerator. Break the cost structure apart: - **The corpus itself: effectively free.** Pairs from other organisations' breaches accumulate, get merged and deduplicated into combined lists numbering in the billions, and circulate at commodity prices or none. Nothing about this input depends on you. - **Address matching: free.** For a consumer login there is no targeting step at all; every address in the corpus is simply tried. - **Source diversity: rented.** Distributed residential source pools exist so that no per-source-address counter at the login accumulates. This is a metered cost and it is the largest line item in most runs. - **Challenge solving: metered per request.** Whatever interactive challenge sits in front of the login is priced as a per-solve fee by third parties who exist for that purpose. - **Compute and bandwidth: negligible** at these volumes. Against that, one million pairs at 0.2% is two thousand valid sessions. Whether the run clears depends only on whether two thousand sessions resell for more than a million metered requests cost — and the per-session resale value of an account with stored payment details, loyalty balance, or simply a reusable identity is far above the per-request price. That is why the run happens at all, and why the operator is entirely content with a rate that would embarrass an engineer. ## The corollary that trips people up Because the input costs nothing, **there is no volume at which the attack becomes uneconomic from your side**. You cannot exhaust the operator's patience or its corpus. Every request you make more expensive raises the price of the marginal run; it never removes the run, because the operator simply raises the price it charges the buyer of the sessions, or shifts to the next target where the same corpus works. ## Which levers touch which term | Change | Term it touches | Effect on the operator | |---|---|---| | Per-source-address throttling | cost per attempt | rents more source addresses; a line item, not a wall | | An interactive challenge at the login | cost per attempt | pays a per-solve fee; a line item | | Longer minimum password length | future hit rate | reduces recoverability from the *next* dump; slow-acting | | Screening at set time against a known-breached corpus | **hit rate now** | removes the intersection between the corpus and your estate | Only the last row changes the numerator. Screening at password-set time, plus a one-off screen of existing passwords with a forced change where they match, takes your accounts out of the corpus. A 0.2% hit rate is 0.2% *because that is the reuse overlap*; drive the overlap toward zero and the run against you returns nothing at any request price. ## Corpus value is not uniform The bulk arithmetic above describes commodity operators working a public consumer login. It is not the whole market, and the contrast is where the economics get interesting. A single reused pair belonging to one open-source maintainer — an address from a decade-old forum dump, matched by hand against a public profile, and a password never changed since — can be worth more than ten thousand consumer pairs. The bulk pairs monetise as interchangeable sessions at a low fixed unit price. The maintainer's pair monetises as reach: publish rights over a package, and thereby a path into every organisation that installs it. The corpus is the same corpus. What changes is that one operator sells volume and the other buys leverage, and the second only needs the reuse to hold **once**. That is why "our users are consumers, the pairs are low value" is a claim about your median account and says nothing about your worst one. ## The reviewer's answer Asked whether a 0.2% conversion means the attack is failing, the answer is: it means the attack is working as designed. The number to ask for is not the hit rate but the overlap — how many accounts in the estate currently hold a password that appears in a known-breached corpus. That number is measurable, it is the operator's actual yield, and screening is the only thing that reduces it.

  • Why can one reused pair belonging to an open-source maintainer outprice ten thousand consumer pairs?
    Because corpus value is not uniform. Bulk pairs monetise as interchangeable sessions at a low fixed price. One reused pair that opens a package-publishing identity monetises as reach over every organisation that installs the package, and matching a decade-old forum dump's address to a public maintainer profile costs a single query. The economics flip from volume to targeting, and the reuse only has to hold once.
  • Does putting an interactive challenge in front of the login remove the attack?
    No. Per-solve services exist precisely to price that away, and the fee is already in the operator's budget alongside proxy rental. A challenge raises cost per attempt, which shifts the marginal run and may push a commodity operator to an easier target, but it leaves the corpus and the reuse overlap untouched. If the pairs still match, the run still works for anyone willing to pay.
  • How would you actually measure your exposure rather than argue about it?
    Measure the overlap: screen the current password of every account against a corpus of known-compromised values and count the matches. That count is the operator's expected yield expressed as a number you own. Repeat it after enforcing screening at set time and forcing a change on matches, and the difference is the only credible evidence the control worked.

saying these in an interview costs you the question

  • Assumes a hit rate under one percent means the run failed
  • Thinks the target must have been breached for stuffing to work
  • Believes per-source rate limiting removes the corpus
  • Prices the attack by hit rate rather than cost per valid session
  • Assumes every stolen pair carries the same value

context