skip to content

You lead a red-team program with in-house GPUs for optimising attack strings against local models, and a set of hosted endpoints you can only send prompts to. Transferred strings tend to stop working after vendor changes. How do you decide how much of a program's effort goes into surrogate-optimised transfer attacks versus attacks discovered by querying the hosted endpoints directly, and what evidence would make you cut transfer work entirely?

level: principalimportance: should knowfreq 28%

answer

  1. measure your own transfer share and survival time
  2. half-life vs. reporting cycle
  3. durable unit = behaviour class + recipe
  4. local search is free of query cost and monitoring
  5. cut when targets are guard-dominated

basics

~20 s

Decide from your own history, not from the literature: track what fraction of locally optimised strings ever fired at a hosted endpoint and how long they kept working. If that half-life is shorter than your reporting cycle, transfer buys exhibits that are stale on delivery. Keep it where it produces cheap candidates; cut it when the endpoints are guard-dominated.

solid answer

~50 s

**Start from the deliverable.** A transferred string is a perishable exhibit. What an owner can act on is a *behaviour class* with evidence and a re-test path, so the allocation argument is about which method produces durable claims per unit of effort. **Measure your own two numbers.** Over past engagements: the share of locally optimised strings that ever produced the behaviour at a hosted endpoint, and how long a working string kept working. A low share means the surrogate guessing is not working; a short survival time means transferred exhibits cannot serve as regression tests. **Where transfer earns its place.** It runs on hardware you own, with no metered queries, so it is a cheap candidate generator and a source of class-level evidence about models you can never instrument. **What would make me cut it.** Sustained near-zero transfer share; targets whose refusals come from surrounding classifiers, so model-level strings never reach the model; or a scope where only the deployed pipeline is actionable.

go deeper

for a junior

Not expected to allocate program effort; should recognise that locally optimised strings may not work on hosted endpoints and do not last.

for a middle

Should compare the two approaches on cost and on what each can actually measure — local search is cheap per attempt, query work measures the deployed system.

for a senior

Should insist on tracking transfer share and survival time from past engagements and let those drive the split, and should write findings as behaviour classes rather than strings.

for a principal

Should treat it as a portfolio decision under a measured decay rate, account for staffing and scarce skills, and state explicit criteria that would retire the capability.

**The question is an allocation under a decay rate, and the decay rate is measurable in your own data.** Almost every argument about this tradeoff is conducted on vibes and on published results from other people's targets. It does not have to be. **Instrument the program before arguing about it.** For every engagement, record four counts and one duration: how many strings were optimised locally; how many produced the behaviour on a held-out local model; how many produced it at a hosted endpoint in scope; how many of those were still working when re-tested; and, from a scheduled re-test cadence, how long each working string survived. After three or four engagements you have an empirical **transfer share** and an empirical **survival time**, and the allocation stops being a matter of taste. **Read those numbers against the reporting cycle.** If working strings survive less time than it takes to write and deliver the report, a transferred exhibit is stale on arrival and certainly cannot serve as the regression test someone re-runs each quarter. That does not make the work worthless; it means the *string* is the wrong unit of delivery. The durable unit is the behaviour class plus the recipe — objective, surrogate choice, ensemble composition, search settings — which lets someone re-derive a candidate against whatever build is live for the price of GPU hours. **Where the numbers mislead, and this is the part that decides budgets.** - *A denominator swap in the transfer share.* Teams quote transfer rates computed only over strings that already passed held-out validation. That figure measures the last stage of the funnel, not the return on the pipeline, and it can look excellent while the end-to-end yield per GPU-week is dismal. Always report the share over strings the search produced, and keep the funnel visible. - *Right-censoring in the survival time.* If you re-test quarterly you cannot observe a three-week half-life; every string that died in week four is recorded as "still working at last check" or as an unexplained gap. A cadence coarser than the decay you are trying to measure guarantees an optimistic answer. - *Spare capacity read as free.* Idle GPUs make the marginal run look costless. The real cost is the people: engineers who can define an objective, pick a surrogate and interpret a failed search are scarce, must be staffed continuously to stay sharp, and are the same people who could be doing query-driven work on the deployed systems that the report is actually about. **What each method uniquely buys.** Local search runs on hardware you own: no per-query cost, no rate limits, no unusual traffic against a vendor's monitoring, and it can run before the engagement window opens. Its output is *candidates*, plus class-level evidence about models you can never instrument. Query-driven discovery is the only thing that measures the **deployed** system — model plus system prompt plus guards — which is normally the thing with an owner who can act. The sane default is a portfolio: local search as a candidate generator, query work as the measurement instrument, split tuned by your measured transfer share rather than by whichever paper was published most recently. **The cut criteria, stated so they can actually fire.** Retire the transfer capability when (a) the measured transfer share stays near zero across several engagements, meaning the surrogate guessing is not working against this target population; (b) the targets you face reject at a surrounding classifier so consistently that model-level artefacts never reach a model — more GPU hours cannot fix a request that is never delivered; or (c) the scope of the work is the deployed product, so a model-level finding has no owner and no remediation path. Even then, keep a small standing capability and the tooling warm: a target population with locally runnable models, or a customer who deploys open weights themselves, brings the whole method straight back, and rebuilding the skill from zero costs far more than maintaining it.

  • What is the durable unit of a finding if the string itself expires within weeks?
    The behaviour class, evidenced by a dated exhibit, plus the recipe — objective, surrogate choice and search settings — that lets someone re-derive a candidate against the current build.
  • Your measured transfer share is decent but survival time is under a month. What changes in how you run the program?
    Keep the local search as a candidate generator, but stop treating strings as deliverables: report classes, attach dated exhibits with raw counts, and schedule re-tests rather than re-quoting the old rate.

saying these in an interview costs you the question

  • Allocating effort from published results rather than from the program's own measured transfer share.
  • Promising transferred strings as durable regression tests.
  • Committing GPU capacity with no plan to measure whether transfer ever pays off.
  • Framing it as an all-or-nothing choice instead of a portfolio with a measured split.
  • Ignoring that the scarce resource is often the people who can run and interpret the search, not the GPUs.

context