skip to content

An adversarial-robustness toolkit's attack classes are written for a particular input data type. What goes wrong when you point an image-oriented gradient attack class at a tabular fraud model whose features include one-hot encoded categoricals, integer counts and an account-age field the customer cannot change?

level: middleimportance: should knowfreq 45%

answer

  1. float array is not a licence
  2. one-hot goes fractional
  3. immutable columns get perturbed
  4. action space before attack class
  5. revalidate rows through the production path

basics

~20 s

It treats every column as a continuous pixel and nudges all of them. You get rows with fractional counts, several one-hot columns partly set so no single category is encoded, values outside legal ranges, and a changed account age. The label flips, but no attacker could submit that row, so the success rate measures nothing actionable.

solid answer

~50 s

Image attack classes assume the input is a dense, continuous, bounded tensor where every coordinate is independently and freely perturbable by a small amount. Tabular data violates all of that. One-hot blocks are mutually constrained; counts are integers; some fields are immutable to the attacker (account age, tenure, verified identity attributes); others are semantically linked (a transaction total and its component amounts). A gradient step ignores every one of those constraints. The visible symptom is a high success rate on rows the production system would reject or that no real customer could produce. The fix is to model the attacker's **action space** first: which columns are mutable, in which direction, in what steps, and under what joint constraints — then use classes that support masking or discrete, constraint-aware transformations, or wrap the model so infeasible candidates are rejected. Text is the same problem in a different costume, which is why text-oriented tooling operates on word- or character-level transformations under explicit constraints rather than adding epsilon to an embedding.

code

text · 2 lines
text
original:   amount=420.00  txn_count_7d=3   acct_age_days=812  cat_online=1.00  cat_instore=0.00
perturbed:  amount=433.71  txn_count_7d=3.47 acct_age_days=809.6 cat_online=0.62  cat_instore=0.38

go deeper

for a junior

Should say the attack produces rows that are not valid inputs — fractional counts, broken one-hot encodings — so the success rate is not meaningful.

for a middle

Explains the assumptions the class carries (continuity, independence, uniform mutability, a meaningful norm) and that mixed-unit tabular data breaks each one.

for a senior

Defines the attacker action space with the domain owner, projects the search onto it, revalidates rows through the production input path, and reports perturbations in domain units.

for a principal

Requires action-space definition as a scoping deliverable for any non-image target, so results are actionable and comparable rather than artefacts of the default class.

### Why nothing errors Access level fails loudly: a gradient class refused a label-only endpoint. The **data type** assumption fails silently, because a tabular row and an image batch are both float arrays. The wrapper accepts the array, the class steps it, the model returns a different label, and the toolkit reports a success rate. Every component behaved correctly. The only broken thing is the meaning of the output. ### The four assumptions an image-oriented class carries 1. **Continuity** — any coordinate can move by an arbitrarily small amount. A pixel can go from 0.412 to 0.417. A transaction count cannot go from 3 to 3.47. 2. **Independence** — coordinates move separately. A one-hot block is a *mutual* constraint: exactly one column is 1 and the rest are 0. A gradient step nudges all of them, and the block stops encoding any category. Derived columns (a total and its parts) are similarly linked. 3. **Uniform mutability** — every coordinate belongs to the attacker. In a fraud model, `amount` might; `acct_age_days`, tenure, KYC-verified attributes and anything the bank computes server-side do not. The attacker cannot age their account backwards. 4. **A meaningful norm** — small distance means imperceptible. On an image this is roughly true. Across mixed-unit standardised columns it is not: an L-infinity step of 0.1 in scaled space might be a few pence on `amount` and a whole extra dependant on a count column, and there is no perceptual system to be fooled anyway. "Imperceptible" is not even the right goal — the goal is "a change the attacker can actually make, cheaply". ### What the run costs, and what it costs you later The compute is trivial; that is part of the trap. A few hundred tabular rows through an iterative gradient class is seconds of CPU, so nobody feels a reason to think first. The real cost lands downstream: an engineer-day or two of triage on findings that turn out to be unsubmittable rows, a meeting with the fraud domain owner to discover that half the perturbed columns are server-computed, and — the expensive one — a remediation programme aimed at a threat that does not exist, because the client hardened a feature no customer controls. Defining the action space up front is typically a couple of hours with the domain owner. It is the cheapest step in the whole exercise and the one most often skipped. ### Where the number misleads The reported attack success rate has an inflated numerator. It counts every flip, including flips produced by rows that the production input path would reject outright (schema validation, range checks, integer types, categorical enums) or — worse — **silently coerce**. Coercion is the subtle case: the pipeline rounds 3.47 to 3 and argmaxes the smeared one-hot back to a single category, which destroys the exact perturbation that flipped the label. The tool's number says 82%; the number an attacker could reproduce through the real intake may be near zero, and nothing in the output distinguishes the two. The perturbation size is misleading in the same way. Reported as an L2 or L-infinity distance in standardised space, it is uninterpretable to the person who has to fix anything. "Norm 0.31" tells a fraud lead nothing; "split the payment into two transfers and declare income 4% higher" tells them everything, including whether their existing controls already catch it. ### Doing it properly, and the residual failure - **Enumerate the action space** with the domain owner: which columns are mutable, in which direction, in what step size, under what joint constraints, and what each unit of change costs the attacker. - **Constrain the search** to that space — a feature mask over immutable columns, an explicit projection back onto the feasible set after each step (round integers, re-argmax one-hot blocks, clip ranges, re-derive dependent columns), or a class built for discrete/constrained domains. - **Re-evaluate the projected row**, not the pre-projection candidate. The flip must survive projection or it is not a flip. - **Revalidate through the production input path.** A candidate that fails the real validator is not a finding. - **Report in domain units** the client can act on. Text is the identical problem in different costume: adding epsilon to an embedding yields a vector with no corresponding string, which is why text-oriented tooling such as TextAttack works through word- and character-level transformations under explicit constraints, so every candidate is a real input the target could receive. Even after all this, a masked and projected row can be **legal but implausible** — a combination no real customer generates, which downstream monitoring would flag on its own. The toolkit has no opinion about plausibility; only the domain owner can rule on it, and that judgement belongs in the finding.

  • How would you constrain the search to the attacker's real action space?
    Mask the immutable columns out of the update, project each step back onto the feasible set (integers rounded, one-hot re-argmaxed, ranges clipped), and re-run the model on the projected row so the reported flip corresponds to a row that could actually be submitted.
  • Why is text a version of the same problem?
    Adding a small perturbation to an embedding produces a vector with no corresponding string. Text tooling therefore works with word- or character-level transformations under explicit constraints, so every candidate is a real input the target could receive.

Perturbing a tabular row with an image attack is like proving a lock is weak by picking it with a key cut in half-millimetre increments: the physics works, but no locksmith can cut 3.47 teeth, so nobody can actually open that door.

saying these in an interview costs you the question

  • Reporting a tabular success rate without checking the adversarial rows are submittable inputs.
  • Treating a small norm in standardised feature space as 'imperceptible' for mixed-unit columns.
  • Perturbing fields the attacker cannot control and counting those flips as successes.
  • Assuming a class runs correctly on tabular data just because it accepted the array without error.

context