skip to content

You are attacking a tabular fraud model with an adversarial-example library, and 12 of its 40 features are ones the attacker cannot influence, such as account tenure. The library lets you pass a mask of which features may move. What does that mask guarantee, and what still has to be enforced outside the library?

level: middleimportance: must knowfreq 60%

answer

  1. mask = which coordinates move
  2. frozen columns bit-identical
  3. no integrality, no one-hot, no derived fields
  4. not every attack consumes it
  5. check the 1-means-movable convention

basics

~20 s

The mask freezes coordinates: features you mark immovable keep their original values, so the search only touches the ones you allow. It says nothing about the movable features' own rules, such as integer counts, one-hot exclusivity, or a field derived from another. Those you check yourself after the library returns its examples.

solid answer

~50 s

The mask is a per-coordinate on/off switch applied to the perturbation, usually by zeroing the update on the frozen positions each step. Its guarantee is exact and narrow: **a masked feature comes back bit-identical to the original**. That correctly models "the attacker controls the transaction fields but not the account history". Three things it does not give you: 1. **No intra-feature type rules.** A movable integer count comes back fractional; a movable categorical comes back as a continuous blend. 2. **No cross-feature rules.** One-hot columns can both go to 0.6; a total that must equal the sum of its parts will not. 3. **No universal support.** Not every attack in a library accepts a mask, and the ones that do differ in shape (per-example versus broadcast). An attack that ignores it will happily move frozen features and give you a higher, wrong success rate. So the mask handles *which* features move; everything about *what values are legal* for the moving ones is a post-generation check you write.

code

python · 3 lines
python
frozen = ~movable            # boolean array over features
assert (x_adv[:, frozen] == x_orig[:, frozen]).all(), \
    "attack moved features the attacker cannot control"

go deeper

for a junior

Knows a mask exists and that it stops chosen features from being modified.

for a middle

Explains it as an elementwise zeroing of the update on frozen coordinates, and lists what it cannot express: types, one-hot groups, derived fields.

for a senior

Asserts frozen columns post-run, checks that the chosen attack really consumes the mask, and pairs the mask with an explicit validity predicate before reporting anything.

for a principal

Decides whether continuous attack libraries are the right instrument for this data at all, versus a discrete search that encodes capability directly.

## Three tiers of constraint, and the library reaches two **Tier 1, the box.** The value range declared on the estimator wrapper (ART's `clip_values`, Foolbox's `bounds`). One global interval, enforced by clipping every iteration. **Tier 2, the mask.** Which coordinates may move at all. In ART this is a `mask` keyword passed to the attack's `generate()` call — `ProjectedGradientDescent.generate(x, y, mask=m)` — where `m` is broadcastable to the shape of `x` and **any position whose mask value is zero is left unperturbed**. Implementation is an elementwise multiply of the update by the mask, so frozen positions accumulate nothing across the whole search. **Tier 3, everything else.** Integrality, categorical one-hot exclusivity, monotone relations between columns, derived fields, checksums, joint plausibility of the record. There is no hook. The attack optimises in continuous space and hands you a float vector. The mask is the tier that encodes **attacker capability**: the transaction fields a fraudster controls versus the account history they do not. That is a genuinely useful thing to state, and it is why a masked run is more honest than an unmasked one. It is not why a masked run is honest. ## What it costs The multiply is free. The interesting cost moves in two directions. Against a white-box gradient attack, freezing 12 of 40 features shrinks the search space, so the attack gets *weaker* and you often need more iterations or restarts to reach a comparable rate — GPU-hours go up for the same claim. Against a black-box attack that estimates gradients by querying, the cost moves the other way: estimation cost scales with the number of coordinates you are estimating over, so masking a third of the features cuts roughly a third of the per-step queries, which on a metered endpoint is money. Either way, a masked and an unmasked run are not the same experiment and their numbers do not belong in the same column. ## Where the number misleads Two specific failures, both of which produce a believable rate. **The mask is silently ignored.** Not every attack class in a library consumes a `mask` argument, and depending on the class and the version an unexpected keyword is either rejected loudly or absorbed into `**kwargs` and dropped. When it is dropped, the run is simply unmasked: the search uses the features the attacker cannot touch, and those are frequently the most predictive ones — account tenure, history length, aggregate counts — so the success rate comes back *higher* and looks like a strong finding. **The mask is inverted.** ART's convention is that a zero means "do not perturb", so a non-zero entry marks a movable feature. Anyone carrying a mental model of a "frozen mask" writes the array the other way round and produces a run that attacked exactly the 12 features the attacker cannot influence and froze the 28 they control. Nothing errors. The number is plausible. It is fiction. And beneath both: even a perfectly honoured mask leaves tier 3 untouched. A fraud finding that rests on moving `prior_disputes` from 2 to 2.37, or on setting two indicators of a one-hot group to 0.6 each, has demonstrated nothing an attacker can submit. The frozen columns being right does not make the moving ones real. ## What I would check Assert, do not assume. After every masked run, compare the frozen columns of the returned examples against the originals elementwise and require exact equality — one line, and it catches both the ignored mask and the inverted one. Then run the attack once with an **all-zero mask**: nothing may move, so success must collapse to the base error rate. If it does not, the argument is being dropped and every masked number in the report is unmasked. Finally, hold a separate validity predicate over tier 3 and report its rejection rate next to the success rate, because that pair is what tells a reader how much of the attack surface the library's relaxation invented. If the attack you need genuinely takes no mask, the cheap fallback is to overwrite the frozen columns from the original row afterwards and re-query the model. Say so in the report: the search was still allowed to use those columns to find its direction, so it is a weaker claim than a truly masked search.

  • How would you handle a one-hot encoded categorical the attacker can change?
    Not through the mask. Let the group move, then project each generated row to the nearest legal one-hot after generation and re-query the model, or search over the discrete category set directly instead of using a continuous attack.
  • The attack silently ignored your mask. How would you find out?
    Compare the frozen columns of the output against the input and assert equality. A run that reports high success while touching frozen columns is the failure this catches.

saying these in an interview costs you the question

  • Assuming a mask makes the generated rows valid records.
  • Passing a mask to an attack that does not accept one and not noticing it was ignored.
  • Never asserting that the frozen columns came back unchanged.
  • Treating a fractional value in an integer-valued feature as a rounding detail rather than an invalid example.

context