skip to content

When does fitting a substitute beat spending queries per input against a rate-limited ranker?

level: seniorimportance: nice to knowfreq 30%

answer

  1. one fixed cost against a recurring one
  2. how many inputs must land
  3. what the endpoint returns sets the per-input price
  4. discount the fixed side by transfer rate
  5. the retrain interval closes the window

basics

~20 s

Fitting a stand-in is one up-front query cost amortised over every later attack; probing per input pays again each time. With many inputs to land, the fixed cost wins — unless the target is retrained before it pays back.

solid answer

~50 s

It is a budget allocation, not a preference between techniques. Attacking the endpoint directly costs queries **per input**: with scores returned you can buy a gradient estimate by probing and differencing them, and with only a verdict you can still walk the boundary from an accepted starting point — but either way the meter runs again for the next creative. Fitting a stand-in pays once, then every later attack is free and offline, at the price of a transfer rate below one. So compare the number of inputs you need to land, times the per-input query cost, against the fixed labelling cost divided by the transfer rate. Two things cap the fixed side: a verdict-only endpoint removes the score-based route entirely and raises the per-input alternative, and the target's retrain interval bounds how long the stand-in stays valid, so the amortisation window is not open-ended.

go deeper

for a junior

Hold on to the shape of it: attacking the endpoint directly costs queries every time you attack a new input, while fitting your own model costs queries once.

for a middle

Be able to state what the endpoint returns and how that sets the per-input price — scores can be differenced, a bare verdict only supports walking the boundary, which needs more queries.

for a senior

Show the arithmetic and both caps: discount the fixed route by transfer rate, and bound its payback by the target's retrain interval rather than assuming a stand-in lasts forever.

for a principal

Own the defensive read: response verbosity and retrain cadence move the attacker's price and the payback window, and should be argued as costs with stated effects rather than sold internally as protection.

## Two ways to buy the same missing information The attacker lacks a direction. There are two ways to pay for it, and they have completely different cost curves. **Per-input spend, against the target itself.** If the endpoint returns a score for the predicted class, an attacker can probe around an input and difference the returned scores to estimate how the score changes with the input — a metered purchase of something that is free when you own the weights. If the endpoint returns only a verdict, that route is gone, but a decision-based approach still works: start from an input the target already accepts and use the returned labels to edge along the boundary toward the input you want accepted. Both spend their queries on **one input at a time**, and both leave a query pattern in the defender's logs tied to that input. **Fixed spend, against a local stand-in.** Buy verdicts on a batch of in-domain inputs the attacker already owns, fit a local model, and from then on optimise offline for free. The queries were spent on *fitting*, not on any particular attack, so they amortise across every creative attacked afterwards. ## The comparison, written out Roughly: attack the target directly when `inputs_to_land x queries_per_input` is smaller than `fitting_queries / transfer_rate`. Fit a stand-in when it is larger. That one line contains the whole judgment, and each term is a real quantity you can estimate before spending anything: - **Inputs to land.** One creative to slip through is a different problem from a campaign of hundreds. This term alone decides most cases. - **Queries per input.** Set by what the endpoint returns. Verdict-only pushes this up sharply relative to a scored endpoint, because boundary-walking is less informative per query than differencing scores. - **Fitting queries.** Small relative to a training set, because the stand-in needs agreement only where the attack operates and unlabeled in-domain inputs are free to the attacker. - **Transfer rate.** The discount on the fixed route. At a 20% transfer rate the attacker must craft five candidates per success — cheap offline, but each replay is still one submission against the rate limit. ## What the rate limit actually does A rate limit does not change which route is cheaper in queries; it converts queries into **calendar time**, and it does so asymmetrically. Per-input attacks consume the limit continuously, for as long as the campaign runs. The fixed route consumes it once, in a burst that looks like ordinary heavy usage, and then goes quiet — the offline optimisation is invisible to the defender entirely. That asymmetry, not the raw query count, is often what decides the attacker's choice, and it is why a defender who watches for long boundary-probing patterns can miss the fitting phase completely. ## The cap nobody remembers: staleness The fixed route assumes the boundary stays where it was. Every retrain, every fine-tune on fresh policy data, every change to the preprocessing chain moves it, and the stand-in degrades. So the amortisation window is bounded by the **retrain interval**, and the real question is not 'is the fixed cost lower' but 'is the fixed cost lower *within one retrain interval*'. A model refreshed weekly forces the attacker to re-pay the fitting cost weekly; a model that has not been retrained in a year gives them an asset that keeps working. This makes retrain cadence a genuine cost control, and it is worth stating its limits honestly: it does not stop the technique, it does not reduce the transfer rate of an attack crafted today, and it costs the defender real engineering effort. It shortens the payback window, which is a different and smaller claim than 'we are protected'. ## Hybrids, and why the split is not clean In practice the routes combine: a stand-in gets the attacker most of the way, and a handful of per-input queries at the target confirm which candidates land. That combination is usually cheaper than either alone, which is another reason to treat 'we removed scores' as a bill rather than a boundary. Hiding scores removes the score-based per-input family and raises the price of the rest; it leaves the fixed route, and the amortised arithmetic, entirely intact.

  • How does hiding confidence scores change this calculation?
    It removes the score-based per-input route, since there is nothing left to difference, and raises the cost of the remaining label-only boundary walk. It does not touch the fixed route: verdicts are still labels, and a stand-in can still be fitted from them. So the effect is to push attackers toward the amortised option — a price change, not a barrier.
  • Why does the target's retrain cadence belong in this decision?
    Because the fixed spend only pays back while the stand-in still agrees with the target. Retraining moves the boundary, so the amortisation window is at most one retrain interval. A weekly refresh forces the fitting cost to be re-paid weekly; a model left untouched for a year hands the attacker a durable asset for one payment.
  • Which route is easier for the defender to notice?
    The per-input route, usually. It produces sustained, structured probing tied to individual inputs over a long period. The fitting phase is a burst of ordinary-looking submissions that stops, after which all optimisation happens offline and is invisible. Detection built only around long probing sequences will miss the amortised route entirely.

saying these in an interview costs you the question

  • Treats the choice as a technique preference, not a budget one
  • Ignores that per-input queries recur for every new input
  • Forgets the transfer rate discounts the fixed route
  • Assumes a fitted substitute stays valid indefinitely
  • Calls hiding scores a boundary rather than a cost increase

context