Why does an API that returns an attribution with each score undercut its own query-budget defence?
answer
- cost per call, value per number
- the limit prices the wrong unit
- d numbers for the price of one
- it still bounds breadth, not depth
basics
~20 sThe price is charged per call, but the value taken is per number returned. An explained reply carries a whole row of local sensitivity instead of one score, so a call limit sized against score-only traffic is loose by roughly the number of features disclosed.
solid answer
~50 sRate limits and paid query budgets are denominated in replies, and that denomination silently assumes a fixed, small amount of model information per reply. That assumption is what the limit actually prices. When a claim-triage endpoint returns a signed contribution for each of, say, forty-seven input fields at the same per-call price as a bare score, the information per reply jumps by roughly that factor — a caller restricted to scores would have to spend many calls around a single record to build a comparable local picture, and here it arrives in one. So a limit calibrated for score-only traffic is now loose by about the disclosure width. What the limit still bounds is **breadth**: how many distinct records the caller can probe at all, and therefore how much of the input space they can cover. It no longer bounds **depth** at each of those records. That is why the limit is not useless — it is mis-denominated.
go deeper
Know that a call limit assumes each reply leaks a small fixed amount, and that returning many numbers per reply breaks that assumption.
Be able to do the arithmetic out loud: identical price per call, many more numbers per call, so the same cap buys roughly the disclosure width times more model information.
Separate what the limit still constrains — how many distinct records get touched — from what it no longer constrains, and say which controls act on which.
Own the process failure: the output contract changed for transparency reasons and nobody re-derived the abuse budget against the new payload or its pricing.
## What a query budget is actually pricing An operator who caps an integrator at some number of calls per day is making an implicit bet: *the amount of model information leaving per reply is small and fixed, so bounding replies bounds the leak.* Every query-budget defence rests on that bet, whether or not anyone writes it down. The bet holds when the endpoint returns a score. It does not hold when the endpoint returns a score plus a per-feature contribution list. ## The arithmetic to internalise Take a claim-triage model over d input fields, exposed to broker integrators, with an identical per-call price whether or not the attribution is requested. - **Score only.** One number per call. To learn how the model responds to each of the d fields near a given claim, a caller has to work for it across many calls concentrated around that one record. Their unit of cost (a call) and their unit of value (a fraction of one local picture) are badly matched, which is exactly what makes the call limit bite. - **Score plus attribution.** One call yields the score *and* d signed magnitudes describing local response. Their unit of cost is unchanged; their unit of value is now a complete local sensitivity row. The daily cap did not move. The value extracted under that cap went up by roughly d. On a forty-seven-field model, a 10,000-call limit that the operator reasoned about as "ten thousand answers" is closer to "ten thousand local response profiles". ## Breadth versus depth — the part candidates miss The limit is not thereby worthless, and saying so is the mark of a shallow answer. - **Depth** — how much the caller learns *at* a record — is what the attribution collapses. It is no longer metered at all, because it arrives whole. - **Breadth** — how many distinct records the caller can touch, and hence how much of the input space they cover — is still bounded by the call limit, and additionally by which claims that integrator can plausibly submit. A broker's coverage is their own book of business; they cannot cheaply obtain claims from segments they do not write. So the correct statement is not "rate limiting is defeated". It is: **the limit is denominated in the wrong unit for this output contract.** It prices replies while the caller now buys sensitivity rows, and it constrains coverage while doing nothing about resolution. ## Where the operator's reasoning went wrong Three failures usually stack up: 1. **The output contract and the abuse control were designed by different people.** The attribution was added for a transparency commitment; the rate limit was sized against traffic patterns that predated it. Nobody re-derived the limit after the payload changed. 2. **Explained calls are not counted separately.** If one counter covers both, the operator cannot even measure how much of their traffic is high-disclosure, let alone price it differently. 3. **Disclosure width was never treated as a dial.** How many fields are named, and to how many decimal places their magnitudes are reported, are choices with a direct multiplier on the leak per call. They usually default to "all of them, full precision" because that is the easiest thing to build. ## What a reviewer concludes The review finding is not "turn off explanations" — the disclosure may be contractually or legally required. It is that the budget must be reasoned about in the unit the caller actually cares about. That points at metering explained calls on their own counter, at deciding how wide and how precise the disclosure needs to be to satisfy the obligation it exists for, and at recognising that whatever number is chosen, it now buys roughly d times what it used to. ## The claim to refuse "We rate limit, so extraction is bounded." A bound exists, but the person quoting it has not multiplied by the disclosure width, and usually cannot say what one call reveals. Ask that question first; the limit is only meaningful once it is answered.
- So is rate limiting pointless for an endpoint that ships explanations?No — it still bounds how many distinct records a caller can touch, and therefore how much of the input space they cover. What it no longer bounds is how much they learn at each of those records. Describe it as mis-denominated rather than defeated; the honest fix is to meter explained calls on their own counter and to re-derive the cap against what one explained call now yields.
- The endpoint charges the same whether or not the attribution is requested. Why does that pricing detail matter?Because it makes the high-disclosure path free relative to the low-disclosure one, so every rational caller takes it and the operator's traffic mix silently shifts to maximum leak. Pricing is one of the few levers that acts on volume rather than on content, and here it was left pointing the wrong way.
- How would you find out how exposed an existing endpoint already is?Ask what one reply contains, not how many replies are allowed: how many fields the contribution list names, at what numeric precision, and whether the same record re-submitted returns stable numbers. Then check whether explained calls are counted separately from plain ones. If nobody can answer, the existing limit was never sized against this payload.
saying these in an interview costs you the question
- Says rate limiting bounds extraction without asking what one reply contains
- Claims the limit is defeated rather than mis-denominated
- Ignores that breadth of input coverage is still constrained
- Never notices the identical price for a much richer reply
- Treats disclosure width and precision as fixed rather than as dials