An adversarial-example library returns 1000 perturbed rows against a tabular loan model and reports that 91% are misclassified, but many rows hold fractional values in integer-only columns and two category indicators set at once. Where do you put the rule the library could not express, and what number do you report instead?
answer
- validate/repair, re-query, then score
- denominator = examples attempted
- report the discard rate too
- reuse the production validator
- repaired row is a new input
basics
~20 sRun the library's output through a validity predicate you write, before you score anything. Keep only rows a real applicant could actually submit, re-query the model on those, and report misclassified-and-valid over examples attempted, together with the share you discarded. The 91% was computed before that filter and overstates what is reachable.
solid answer
~60 sThe rule lives in a **post-generation validation stage you own**, between the library's output and your scoring. Concretely: 1. Write a predicate over a row: types, integrality, one-hot exclusivity, derived fields, allowed value sets, any cross-field invariant. 2. Apply it to the returned examples. Optionally *repair* instead of rejecting — round, snap each one-hot group to its argmax, recompute derived fields. 3. **Re-query the model on the surviving or repaired rows.** This step is the one people skip. A repaired row is a different input, and rounding frequently pushes it back across the decision boundary. 4. Report the rate as valid-and-still-misclassified over all examples attempted, with the rejection or repair-survival share stated next to it. The denominator matters as much as the numerator: scoring survivors over survivors turns a filter into a way to make any number look good. The pre-filter figure is not wrong, it just measures the wrong space — it is the rate of an attack that was allowed to write inputs the system cannot receive.
code
python · 5 linesx_valid = repair(x_adv) # round ints, argmax one-hots, recompute derived
ok = is_constructible(x_valid) # your predicate, ideally the production validator
preds = model.predict(x_valid[ok])
hits = (preds != y_true[ok]).sum()
print(f"attempted={len(x_adv)} valid={ok.sum()} valid_hits={hits}")go deeper
Recognises that fractional counts and double-set categories mean the rows are not real, and that some filtering is needed.
Describes the validate-or-repair stage and knows the model must be re-queried on repaired rows.
Fixes the denominator, reports counts and discard rate, reuses the production validator, and states that the filtered number is a lower bound.
Decides what claim the engagement is allowed to make from a relaxed-then-filtered run, and whether the budget goes to a constrained search instead.
## Why the stage exists at all The library reaches a value box and a movable-feature mask. A rule past those — an integer count, a one-hot group, a field derived from two others, a checksum, a joint-plausibility rule — has nowhere to live inside the attack, so it lives in a stage you own between the library's output and your scoring. The pipeline is therefore four steps, not two: **generate, validate or repair, re-query, score**. **Validate** means a predicate over one row: types, integrality, allowed value sets, exclusivity within each one-hot group, recomputed derived fields, cross-field invariants. **Repair** means snapping the row to the nearest legal point instead of discarding it — rounding integer columns, taking the argmax of each one-hot group, recomputing anything derived. **Re-query** is the step people skip, and skipping it is what turns a report into fiction: a repaired row is a *different input*, and the model has not been asked about it. ## What it costs The predicate is hours of engineering if you write it, and close to zero if you reuse the production validator, which is also the more defensible choice — a red-team re-implementation is almost always more permissive than the real intake path, and every extra row it admits inflates the result. Re-querying costs one model call per surviving or repaired example: free on a local model, real money on a metered endpoint, and it is the reason people quietly keep the stale labels. Against the 1,000-row run in the question that is at most 1,000 additional calls, which is a trivially cheap way to avoid publishing a number that does not survive its first challenge. ## Where the number misleads **The 91% is the rate of an attack allowed to write inputs the system cannot receive.** It is not wrong arithmetic; it measures the wrong space. Fractional counts and two simultaneously-set category indicators are not near-misses, they are rows no intake path would accept. **Survivors over survivors.** The tempting fix is to filter, then divide the misclassified survivors by the survivors. That makes the filter incapable of lowering the number — the tighter your predicate, the better your attack looks, which is the signature of a broken denominator. Three denominators are defensible and you must name which you used: valid-and-misclassified over **examples attempted** (the attacker's-eye view, and the one I report); over **originally correctly classified examples** (the usual convention when the counterpart figure is robust accuracy); over **valid examples produced** (almost never what a reader assumes). Publish the counts, not only the percentage: attempted, valid, valid-and-still-misclassified. **Stale labels after repair.** Adversarial examples sit just past a decision boundary. Rounding is a move the attack did not optimise for, and it routinely pushes the row back to the correct class. Keeping the pre-repair success label is the most common single defect in tabular robustness reporting. **The filtered number is still only a lower bound.** It came from an unconstrained search whose illegal results you threw away, so the survivors are the ones that happened to land legally. A low filtered rate does not license "the model is robust"; a search that respects the constraints while optimising can beat it. ## What I would check That the predicate is literally the production validator, or a documented subset of it, rather than a parallel re-implementation. That perturbation size is measured in the same space the validation runs in — if the attack worked on scaled features and the validator works on raw ones, both the distance and the legality claim are about different objects. That every repaired row was re-queried. And that the discard rate appears in the report next to the success rate, because a 91% raw figure against a low double-digit valid yield is itself the headline finding: the reachable attack surface here is a small fraction of what the library's output suggests.
- Why must you re-query the model after rounding an integer column?Because rounding is a perturbation the attack did not choose. Adversarial examples often sit just past the boundary, and snapping to the nearest legal value routinely pushes them back to the correct class.
- Is a low rate after filtering evidence that the model is robust?No. It bounds what an unconstrained search that got lucky achieved. A search that enforces the constraints while optimising can do better, so it is a lower bound only.
Scoring the hits only against the arrows that stayed on the board makes every archer look accurate. The denominator has to stay at the number of arrows fired.
saying these in an interview costs you the question
- Reporting the library's raw success rate for tabular data with no validity filter at all.
- Repairing rows and keeping the original success labels without re-querying the model.
- Scoring survivors over survivors so the filter can only improve the number.
- Writing a looser validity predicate than the production system's own input validation.