skip to content

Signal per Reply

A hard label, a full score vector or a returned attribution are three different gifts to a paying attacker. Interviewers ask because output design decides what a copy gets built from.

on this pageshow

explore

questions

8

A claim-triage API returns a per-feature contribution list with each score. What does that hand the caller?

level: juniorimportance: must knowfreq 50%

answer

  1. one number versus a row of numbers
  2. it describes response, not the record
  3. sensitivity belongs to the model
  4. black-box caller, locally white-box answer

basics

~20 s

It hands the caller the model's local sensitivity: for that record, which input fields move the score, in which direction, and roughly how much. That is information about the model, not only about one decision.

solid answer

~40 s

A score is one number about one record. A per-feature contribution list is a set of numbers about how the model *responds* around that record: raise this field and the score goes up by roughly this much relative to the others. That is local sensitivity, and sensitivity is a property of the model, not of the claim. A caller with no other access is normally in a black-box position — they see outputs and nothing else, and any picture of the model's local response has to be inferred from many replies. Shipping the attribution moves that caller a long way toward a locally white-box position for free. The usual wrong answer is that explanations describe the decision rather than exposing the model; a local description of the model's response is model information by definition.

go deeper

for a junior

Be ready to say plainly what arrives in each reply: one score is a fact about one record, while a per-field contribution list is a statement about how the model responds. Do not stop at calling it transparency.

for a middle

An interviewer expects you to place the attribution on the access ladder — outputs only, scores, gradients — and explain that it delivers approximate local sensitivity to someone who otherwise sees only outputs.

for a senior

Show that you treat every party holding an integration key as receiving this disclosure, and that you can say what it buys them and what it still does not, without overstating either.

for a principal

Own the fact that the transparency obligation and the extraction exposure are the same disclosure, so this is a scoping decision about who gets what, not a bug to be patched.

## What the endpoint sells and what it gives away Consider a commercial-lines insurance claim-triage model exposed to broker integrators. A broker submits a claim; the endpoint returns a referral score — how strongly this claim should be pulled out for human review — and, by contract, a per-feature contribution list: which submitted fields pushed the score up, which pushed it down, and by how much. Those are two very different objects. - **The score** is one number about one record. It says what the model concluded. - **The contribution list** is a statement about how the model *behaves near* that record: if this field were somewhat different, the score would move in this direction by roughly this amount, relative to the other fields. The second is **local sensitivity**. Sensitivity is a property of the fitted model, not of the claim that was submitted. The same claim scored by a different model would come back with a different contribution list; that is precisely because the list describes the model. ## Why "it only explains the decision" is the wrong answer The intuition behind that answer is that an explanation is *about the outcome* — a justification, a reason a human can read — and that reasons are downstream of the model rather than part of it. The intuition breaks because the only thing that can justify one prediction is a statement about what the model does with the inputs. There is no way to say "your prior-claims count is what drove this" without also saying "this model is strongly responsive to prior-claims count in this region of the input space". Transparency and disclosure of model behaviour are the same act. ## The access ladder Adversaries against a deployed model are usually placed on a ladder of what they can see: | Vantage | What arrives per query | |---|---| | Weights and gradients in hand | everything; the exact response to any input change | | Full probability vector | the model's confidence over every class | | Score for the predicted class | one number | | Top-1 label only | one symbol | A returned attribution does not sit cleanly on that ladder — it is a *different* gift. It is not the exact input gradient: it may be produced by any of several attribution methods, those methods disagree with each other, and how faithfully any of them describes the underlying model is a live question in the interpretability literature. But for a caller who wants to know which direction to push, an approximate, signed, ranked answer delivered with no work is worth far more than an exact one they would have to pay to reconstruct. **The attacker needs a direction, not a proof.** ## What that opens Three things change for a caller who receives attributions: 1. **Supervision per query.** Each reply is no longer one labelled example for a stand-in model; it is a labelled example plus a description of how the target responds around it. Fitting a functional copy against outputs *and* local responses is a much richer signal per paid call than outputs alone. 2. **A search direction, free.** Where the model is fragile near this record is stated outright rather than inferred. 3. **A map of the business policy.** Which fields dominate a referral decision is commercially sensitive on its own, independent of any copy or evasion. ## What it does not open Be precise here, because overstating is the other failure mode. - It does not give the **parameters** or the architecture. - It says nothing about the model's behaviour on records the caller cannot legitimately submit — their coverage is bounded by their own book of business. - It is **local**. Sensitivity at one record is not a global description, and extrapolating it far from that record is where it stops holding. - It is not, by itself, a flipped decision. Many of the fields in a claim are attested facts rather than free variables, and knowing which direction helps is a long way from being able to move in it. ## The tension worth naming An explainability requirement and an extraction defence pull in opposite directions. The requirement usually exists for a good reason — a decision subject is owed an account of a decision made about them. The extraction exposure comes from the same disclosure being available, at machine rate, to a party who is not the decision subject. Recognising that the two goals genuinely conflict, rather than pretending the explanation is inert, is what the question is testing.

  • Is the attribution the same thing as the model's input gradient?
    Not exactly. Attribution methods differ from each other and from the raw derivative, and how faithfully any of them describes the model is a real open question. But an adversary is after a signed, ranked indication of which fields move the score and which way. An approximation delivered for free is more useful to them than an exact quantity they would have to buy.
  • Does it matter that the caller is a paying, contracted integrator rather than an anonymous attacker?
    It changes who you can sue, not what they receive. The contract entitles them to the disclosure, so every call they make is legitimate on its face, and there is no per-call signal separating business use from collection. Treat the disclosure as public to every party holding an integration key.
  • The vendor says the explanation is post-hoc and approximate, so nothing about the model leaks. Rebut it.
    Approximate information about the model is still information about the model. The approximation degrades a copyist's fidelity a little; it does not change the sign or the ranking of which fields drive the score, which is the part with attack value. Faithfulness is an argument about the explanation's honesty toward users, not about its inertness toward adversaries.

A price is one fact about one purchase. A price list showing how the price changes with every option is a fact about the seller.

saying these in an interview costs you the question

  • Claims explanations describe the decision, not the model
  • Says post-hoc attributions are approximate so nothing leaks
  • Treats the score and the contribution list as equivalent disclosures
  • Assumes a contracted integrator is not an adversary
  • Thinks local sensitivity describes the record rather than the model

context

open as a page

Why does normal paid use of a content-moderation API also hand the caller a labelled dataset?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Because the caller supplies the text and the endpoint supplies the annotation. Every paid call returns a scored row they can keep, so ordinary integration quietly accumulates the labelled corpus your annotators were paid to produce.

open as a page

Your moderation API owner says full category scores are safe since the weights stay private. What's wrong?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Nobody fitting a copy wants the parameters. A returned score vector is already shaped like a training target, so the endpoint acts as a teacher and a competitor buys a functional stand-in at query prices.

open as a page

Why does an API that returns an attribution with each score undercut its own query-budget defence?

level: middleimportance: should knowfreq 40%

basics

~20 s

The price is charged per call, but the value taken is per number returned. An explained reply carries a whole row of local sensitivity instead of one score, so a call limit sized against score-only traffic is loose by roughly the number of features disclosed.

open as a page

In a multi-label moderation API, how does the response body set an extraction campaign's cost per labelled row?

level: middleimportance: should knowfreq 52%

basics

~20 s

One call returns one score per policy category, so a single payment annotates that text against every category at once. The response schema, not the call count, fixes how much supervision a dollar of API spend buys an attacker.

open as a page

A broker integrator archived a year of explained claim-triage replies. What can they build, and what stays out of reach?

level: seniorimportance: should knowfreq 33%

basics

~20 s

They hold claims paired with scores and per-field sensitivity rows — enough to fit a close functional stand-in over the segment their own book covers, and to read off which fields drive referrals. Not the parameters, not behaviour outside that segment, and not yet a changed decision.

open as a page

Transparency commitments force per-field reasons on every model decision. What do you actually negotiate?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Not whether to explain, but the shape of the disclosure: how many fields it names, at what numeric precision, to which recipient, at what rate, and at what price. The obligation usually runs to the decision subject, one record at a time, not to a machine integrator pulling thousands a day.

open as a page

How do you price the per-category scores in a moderation API's response at a design review?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Report what the field buys an adversary in money and time: the annotation budget a competitor no longer funds, and how fast a stand-in reaches usable quality. Leave the keep-or-cut call to the product owner.

open as a page