skip to content

In TextAttack, an attack recipe pairs a transformation that proposes candidate rewrites with a set of constraints those candidates must pass. What role do the constraints play in the search, and what happens to your reported success rate if you relax them?

level: middleimportance: should knowfreq 45%

answer

  1. transformation proposes, constraints prune
  2. text's stand-in for range and mask
  3. constraint stack can dominate runtime
  4. looser constraints, inflated success
  5. label may be stale if meaning moved

basics

~20 s

Constraints are filters inside the loop: a transformation proposes candidate sentences, each constraint rejects the ones that violate it, and the search only ever sees the survivors. They are what keeps a rewrite readable and meaning-preserving. Loosen them and the success rate rises, because you are now counting rewrites that changed the sentence.

solid answer

~60 s

The recipe is a composition: something proposes candidate edits, constraints filter them, a search strategy picks among the survivors, and a goal function says when the model has been flipped. The constraints are the **domain rule for text** — the equivalent of a value range and a mask in the numeric libraries, except that here they sit inside the loop by design rather than being something you bolt on afterwards. Two consequences interviewers look for: - **Cost.** Every constraint runs on every candidate, and the semantic ones are themselves model calls. On a wide candidate set the constraint stack, not the target model, dominates wall-clock time. Ordering cheap checks before expensive ones matters. - **Comparability.** The success rate is defined relative to the constraint set. Loosening a similarity threshold or allowing more of the sentence to change raises success while quietly changing what "success" means — the flipped label may now be the *correct* label for a sentence whose meaning moved. A text attack number is not reportable without stating the constraints it ran under.

go deeper

for a junior

Knows constraints exist to stop the rewrite becoming nonsense or changing the meaning.

for a middle

Places them as an in-loop filter over candidates, and explains that relaxing them raises success by widening the feasible set.

for a senior

Profiles the constraint stack as a runtime cost, reports the thresholds with the rate, and samples successful rewrites for stale labels.

for a principal

Sets the house convention for which constraint stack results are reported under, so numbers across teams and models remain comparable.

## The composition, named Text has no continuous box to clip into, so TextAttack replaces the value range and the movable-feature mask with an explicit four-part object. A `textattack.Attack` is built from a **transformation** (proposes candidate rewrites, for example word substitutions at chosen positions), a list of **constraints** (rejects candidates), a **search method** (decides which surviving candidate to take next), and a **goal function** (queries the target model and says when the attack has succeeded). A published recipe is nothing more than a specific choice of those four, and the CLI runs one with `textattack attack --recipe <name> --num-examples 100`. Constraints come in recognisable families. *Preservation* constraints ask whether the rewrite still means and reads like the original — `WordEmbeddingDistance` on the substituted word, `UniversalSentenceEncoder` on the whole sentence, a part-of-speech check on the swap. *Capability* constraints limit how much may change at all — a cap on the fraction of words perturbed, a rule forbidding repeated modification of the same position or modification of stopwords. Both families are **your assumption about the attacker and the domain**, not a fact about either. ## Where they sit, and why that is the good news Constraints gate the candidate set at each step of the search, before the goal function spends target-model queries. That is structurally better than the numeric libraries, where the equivalent rule is code you bolt on after generation: here the feasible set is declared, inspectable and inside the loop. It is also the reason a TextAttack result is reportable at all — you can state exactly what capability and what preservation rule you assumed. ## What it costs The cost profile is the opposite of what people expect. The naive assumption is that the target model dominates. In practice a transformation can propose hundreds of candidates per position, and a semantic-similarity constraint is itself an encoder forward pass **per candidate**, so the constraint stack routinely dominates wall-clock time — a run that is mysteriously slow against a small fast local model almost always has an expensive constraint evaluating a very wide candidate set. The levers are ordering cheap constraints before expensive ones so most candidates die early, capping the candidates considered per step, caching encoder outputs, and setting an explicit query budget so a single stubborn example cannot consume the run. Budget arithmetic worth doing before you launch: candidates per position times positions times examples is the constraint-evaluation count, and it is usually one to two orders of magnitude larger than the target-model query count. ## Where the number misleads **Success is defined relative to the constraint set.** Loosening a similarity threshold reliably raises the reported rate, and part of that rise is not adversarial at all: the rewrite genuinely changed the sentence, so the model's new prediction may be *correct* and the original label is stale. That is exactly the invalid-row failure from the tabular case arriving through a different door. Two runs under different constraint stacks measure different things and must never be presented as a before-and-after about the model. **The denominator hides a second trap.** TextAttack's summary reports successful, failed and *skipped* attacks, where skipped means the model already got the example wrong, and the attack success rate is computed over the attacked examples only. A weak model skips a large share of the set, so its success rate is computed over the small favourable subset it originally got right. Comparing two models on that figure is comparing two different denominators; **accuracy under attack**, reported in the same summary, is the figure that stays comparable. ## What I would check The per-constraint rejection rate: a constraint that rejects essentially every candidate means the recipe is doing no search and the low success rate is an artefact of your own filter. Where the wall-clock actually went, so the cost story in the report is true. The skipped count alongside the success rate, and original accuracy beside accuracy under attack. And a human read of a sample of successful rewrites — the only real check that the meaning survived, and the one that catches a stale-label inflation no metric will show you.

  • Why can loosening a semantic-similarity constraint produce successes that are not really attacks?
    Because the rewrite may genuinely change the meaning. The model's new prediction can then be correct for the new sentence, so the flip counted as a hit against a stale label.
  • A run against a small local model is far slower than expected. Where do you look first?
    At the constraint stack. Semantic constraints are model calls run per candidate, and with a wide candidate set they can cost more than querying the target.

saying these in an interview costs you the question

  • Reporting a text attack success rate without stating the constraint set it ran under.
  • Comparing two runs with different constraint thresholds as if the difference were about the model.
  • Assuming the target model is always the runtime bottleneck.
  • Never eyeballing successful rewrites to check the meaning survived.

context