skip to content

You lead ML security for several product teams and must standardise how adversarial-robustness evaluations are run. How do you decide between adopting a third-party attack harness, building a thin in-house driver over the attack libraries, or letting each team script directly — and how do you keep that decision reversible?

level: principalimportance: should knowfreq 30%

answer

  1. standardise contract and schema, not the tool
  2. harness for breadth, library for claims
  3. thin driver stays thin
  4. disposable envs versioned with results
  5. leaked tool vocabulary is a chain

basics

~20 s

Decide by what must be repeatable across teams versus what each team must tune. Adopt a third-party harness for a shared target contract and result format; keep the attack calls in library code you own so parameters stay reachable. Keep it reversible: the harness should be swappable without rewriting past evaluations.

solid answer

~50 s

Separate two things that get conflated. **Standardisation** is about the target contract, the result schema and the reporting bar — those must be common, or nobody's numbers compare. **Attack execution** is about tuning per model and per threat model, and must stay flexible. Given that split: - *Third-party harness* fits when teams are many and models conventional; its parameter surface and pins become your ceiling. - *Thin in-house driver* fits when models are heterogeneous, when you need full parameter reach, or when the third-party pins do not fit the estate; you then maintain it. - *Each team scripts directly* is defensible only very early, and costs comparability immediately. Reversibility comes from owning the interfaces rather than the tool: define the result schema and target contract yourself and let any harness fill them. Swapping tools then does not invalidate history.

go deeper

for a junior

Can state that a shared tool makes results comparable and that teams scripting separately produce numbers that cannot be compared.

for a middle

Weighs adopt versus build on parameter reach, pins and estate fit, and suggests a hybrid of harness for sweeps and library calls for reportable numbers.

for a senior

Designs the target contract and result schema first, isolates evaluation environments, and treats the tool as swappable underneath a stable reporting bar.

for a principal

Owns the standard rather than the tool, engineers reversibility through schema ownership and retained raw artefacts, and names the signals that the choice went wrong.

This is a platform decision wearing a tooling costume, and strong answers treat it that way. The question is not "which harness"; it is "what must be identical across teams for their numbers to sit in one table". **Frame it as: what must be common?** Three artefacts, none of which is a tool. 1. *The target contract* — how a team describes the thing being evaluated: the model interface, the access mode (white-box gradients, score-only, decision-only), the preprocessing applied before the model sees an input, and the input range the perturbation budget is expressed in. Without a shared contract, an `eps` from team A and an `eps` from team B are different physical quantities. 2. *The result schema* — what a run records: library and version, attack, every setting, the query or compute budget, the access assumption, the denominator, the success criterion, and the outcome. This artefact outlives every tool change. 3. *The reporting bar* — what an evaluation must contain before a claim may be drawn from it: a swept budget rather than a single point, a sanity check against an undefended reference model, and full provenance. Standardise those and the harness becomes an implementation detail rather than the contract. **Then choose the layer.** *Adopt* a third-party harness when its covered input domains and access modes match most of the estate, its pinned stack can live in an isolated image, and its surfaced parameters suffice for triage — adoption is near-free and the maintenance is someone else's, provided upstream is still alive; check the commit history, because a dormant wrapper hands you the maintenance anyway, later and by surprise. *Build a thin in-house driver* when you need parameter reach the harness does not offer, when the estate is unusual, or when the pins cannot be made to fit; keep it genuinely thin — it calls attack libraries, it does not reimplement attacks — and let it die when a third-party option becomes adequate. Budget it honestly: a thin driver is a couple of engineer-weeks to write and a standing tax of a few days a quarter chasing upstream API changes. *Let teams script directly* only while there are one or two teams and no cross-team claim is being made; it costs comparability from the first week. The usual good answer is a hybrid: harness for breadth-first sweeps and routine regression, direct library calls for anything that becomes a published claim. **Where the number misleads at this altitude.** A cross-team robustness dashboard is the most dangerous artefact this decision produces, because it makes non-comparable numbers look ranked. A team can appear robust because its harness surfaced fewer parameters, ran a smaller attack menu, used the all-inputs denominator rather than originally-correct, tested at a smaller budget, or hit a query cap against a metered endpoint and truncated the sweep. Every one of those reads on the tile as strength. Before any ranking is published, the dashboard must carry the budget, the denominator, the access mode and the attack set per row — and a row that cannot state them shows as "not measured", not as a good score. The corollary: never let "the tool doesn't support that" quietly become "that risk is fine". Record it as unmeasured. **Keeping it reversible.** Reversibility is engineered, not hoped for. Own the schema and make every tool a producer that fills it. Keep evaluation environments disposable and versioned with results, so a tool change is a new image benchmarked against the old rather than a mutation. Retain raw artefacts — original inputs, perturbed inputs, model decisions — beside the summaries, because those are what a future tool can re-score; a rendered summary table cannot be re-scored by anything. And keep the tool's bespoke vocabulary out of your reporting language: every leaked concept is a chain to that tool. **How to tell you chose wrong.** Teams routing around the standard with private scripts; quarter-over-quarter numbers that cannot be compared because the stack moved and nobody froze it; and the phrase "our tool doesn't support that" appearing as the reason a risk went unmeasured rather than as an entry on the risk register.

  • What single artefact makes the harness choice reversible?
    A result schema you own — library, attack, settings, budget, access assumptions, success criterion, outcome — that any tool can fill. History stays comparable when the tool changes.
  • When is a thin in-house driver the wrong call?
    When the estate is conventional and a third-party harness already covers it. You then pay maintenance for a layer that buys nothing, and it will lag the libraries just as a third-party one would.
  • How do you keep a harness from becoming the ceiling on what gets measured?
    Make direct library calls a normal, sanctioned path, and require that any risk skipped because of tooling limits is recorded as such rather than silently dropped.

saying these in an interview costs you the question

  • Picking a tool first and letting its concepts become the reporting standard.
  • Standardising attack parameters across heterogeneous models instead of standardising the contract and schema.
  • Building an in-house driver that reimplements attacks rather than calling the libraries.
  • No plan for what happens when the adopted harness stops fitting the estate.
  • Accepting 'the tool doesn't support it' as a reason a risk was not measured.

context