Who is the plausible adversary against a demand forecaster that reorders stock from supplier-submitted documents?
answer
- start from who can write the inputs
- not everyone can reach this model
- a few fields on a business document
- no queries, no scores coming back
- and would a wrong order even pay them
basics
~10 sOnly a party who can write into the model's inputs. Here that is a supplier authoring lead-time, pack-size and price fields on documents the forecaster reads, with no query access and no scores returned.
solid answer
~50 sStart from who writes each input rather than from a list of attack names. A demand forecaster that turns sales history, stock levels and supplier-submitted trade documents into automatic reorder quantities has exactly one writer outside the trust boundary: the supplier, on the handful of fields they legitimately author. That is their whole vantage. They cannot query the forecaster, there is no endpoint returning a score or a probability, and they see nothing of its parameters. The only feedback they get is the purchase order that eventually arrives, which is one coarse, delayed observation per ordering cycle. So the adversary is real but slow and blind, and the second half of the question is whether a wrong reorder quantity is worth anything to them once a buyer approves high-value orders and every automatic order stays cancellable for a day.
go deeper
Be ready to answer 'who could attack this model' by listing who writes each input and which of them sits outside the company, rather than by naming attack families you have read about.
Explain what the writer can and cannot do: which fields they set, whether anything comes back to them, and why no returned score means no cheap search for an effective change.
Show that you finish the analysis: a writer is only an adversary if a wrong output pays them and survives whatever human step and reversal window sits after the model.
Own the framing that 'no motivated adversary here' is a legitimate, defensible finding, and that recording who the writers are is what makes the finding re-checkable when the product changes.
## Ask who can write, not what the attack is called The reflex when someone says "adversarial machine learning" is to reach for a family of attacks. The useful first move is the opposite: enumerate every party who can put a value into this model, and ask which of them sits outside your trust boundary and gains something from a wrong output. Take a concrete deployment. A retailer runs a demand forecaster over time series: it consumes internal sales history, current stock levels, a calendar of promotions, and trade documents submitted by suppliers carrying lead times, minimum order quantities, pack sizes and price breaks. Its output is a reorder quantity. Below a value threshold that quantity becomes a purchase order with no human involved; above it, a buyer approves. Every automatic order is cancellable for a stated window. ## The writer inventory - Sales history: written by the retailer's own point-of-sale systems. Inside the boundary. - Stock levels: written by the retailer's warehouse systems. Inside. - Promotion calendar: written by the retailer's own merchandising team. Inside. - Supplier trade documents: written by a counterparty. **Outside.** So the entire external adversary surface of this model is a small set of fields on documents a named counterparty legitimately authors. Whether those fields only move the current forecast or also land in a later refit, the vantage is the same set of fields. ## What that vantage actually permits The adversary can choose the values of those fields, inside what a plausible trade document allows and what their contract permits, across the documents they submit over time. What they cannot do is as important: - **No queries.** There is no endpoint they can hit repeatedly with candidate inputs. - **No scores.** Nothing returns a probability, a confidence or a per-feature explanation, so there is no quantity to difference and no estimated direction to follow. - **No parameters.** They hold no weights and no gradients, and no released copy of the model exists. - **Attributable.** Every field they write is signed by their own account, on their own document, under a commercial contract. The one channel back is the purchase order that eventually arrives. It is an oracle of a kind, but a very poor one: roughly one observation per ordering cycle, delayed by days or weeks, and confounded by everything else that moves demand. That matters because most of the adversarial-ML literature assumes fast, cheap feedback. An adversary who can submit inputs and read scores can buy an estimate of the direction that moves the output, by probing and differencing what comes back. An adversary who sees only a returned decision can still walk along the boundary from a point the model already accepts, given enough replies. Both families are paid for in queries. Remove the replies and both are gone; what remains is blind, plausible-looking manipulation with a feedback loop measured in weeks. ## The second half: would it pay? Naming a writer is not naming an adversary. The supplier profits only if an over-sized order becomes money, and three facts stand between them and that: 1. A value threshold routes any large order to a buyer who looks at it. 2. A cancellation window lets a noticed order be reversed before it is a commitment. 3. The fields that caused it are attributable to them, and inflating a lead time to induce over-ordering is a contract problem, not just a model problem. Against the alternative of simply negotiating a larger order, the attack is slow, risky and traceable. For many internal models the honest conclusion is that no motivated adaptive adversary exists at all, and that conclusion is a finding, not an omission. ## Answers that fail - **"Any internet user."** The model has no public interface. Its writers are named counterparties. - **"Nobody, it is internal."** Wrong in the other direction: the supplier document path crosses a trust boundary, and "internal" describes the hosting, not the data. - **"They would need gradients, so it is impossible."** Influence over the values a model reads or refits on needs no gradients at all. Gradients buy efficiency, not access. - **"So we should adversarially train it."** That is a purchase, and the point of the writer inventory is to decide whether it is worth making here. ## The check, as a short list Who writes each input; is that writer attributable; what does the model's output do with no human in the path; is that action reversible; and does a wrong output pay whoever wrote the input. The machine-learning-specific part is the first item: for a model, inputs are data that shapes behaviour, not requests that are merely served, so a party who submits routine business documents is a writer whether or not anyone thought of them as a user.
- The forecaster also ingests a public commodity price feed. Does that change your answer?Yes. Anyone who can move or spoof that feed becomes a writer, and unlike the supplier they may be unattributable and unbounded by a contract. The feed's provenance becomes the control that matters: who signs it, whether values are cross-checked against a second source, and whether an implausible jump is rejected rather than forecast on.
- Why does having no scores returned matter so much to what this supplier can do?Because search needs feedback. With returned scores an adversary can probe and difference them to estimate which direction moves the output; with only decisions they can still walk a boundary given many replies. Both are paid for in queries. With no replies at all, the supplier is reduced to blind, plausible values and one delayed observation per ordering cycle.
- Does concluding there is no motivated adversary mean the model needs no work?No. It means robustness against a deliberate perturbation is not the work. The same wrong reorder quantity can come from a broken feed, a unit mix-up or a genuine demand spike, and those need input validation, sanity bounds and a reversible decision. The finding narrows what you buy; it does not empty the budget.
It is the difference between asking who could pick this lock and asking who has ever been handed a key. Only the second list is short enough to act on.
saying these in an interview costs you the question
- Says any internet user can attack an internal model
- Assumes an adversary exists before listing who writes inputs
- Treats every input source as equally trusted
- Claims an attack is impossible without gradients
- Confuses noisy supplier data with a deliberate adversary
- Names attack families instead of naming a writer