You lead a red-team program with in-house GPUs for optimising attack strings against local models, and a set of hosted endpoints you can only send prompts to. Transferred strings tend to stop working after vendor changes. How do you decide how much of a program's effort goes into surrogate-optimised transfer attacks versus attacks discovered by querying the hosted endpoints directly, and what evidence would make you cut transfer work entirely?
answer
- measure your own transfer share and survival time
- half-life vs. reporting cycle
- durable unit = behaviour class + recipe
- local search is free of query cost and monitoring
- cut when targets are guard-dominated
basics
~20 sDecide from your own history, not from the literature: track what fraction of locally optimised strings ever fired at a hosted endpoint and how long they kept working. If that half-life is shorter than your reporting cycle, transfer buys exhibits that are stale on delivery. Keep it where it produces cheap candidates; cut it when the endpoints are guard-dominated.
solid answer
~50 s**Start from the deliverable.** A transferred string is a perishable exhibit. What an owner can act on is a *behaviour class* with evidence and a re-test path, so the allocation argument is about which method produces durable claims per unit of effort. **Measure your own two numbers.** Over past engagements: the share of locally optimised strings that ever produced the behaviour at a hosted endpoint, and how long a working string kept working. A low share means the surrogate guessing is not working; a short survival time means transferred exhibits cannot serve as regression tests. **Where transfer earns its place.** It runs on hardware you own, with no metered queries, so it is a cheap candidate generator and a source of class-level evidence about models you can never instrument. **What would make me cut it.** Sustained near-zero transfer share; targets whose refusals come from surrounding classifiers, so model-level strings never reach the model; or a scope where only the deployed pipeline is actionable.
go deeper
Not expected to allocate program effort; should recognise that locally optimised strings may not work on hosted endpoints and do not last.
Should compare the two approaches on cost and on what each can actually measure — local search is cheap per attempt, query work measures the deployed system.
Should insist on tracking transfer share and survival time from past engagements and let those drive the split, and should write findings as behaviour classes rather than strings.
Should treat it as a portfolio decision under a measured decay rate, account for staffing and scarce skills, and state explicit criteria that would retire the capability.
**The question is an allocation under a decay rate, and the decay rate is measurable in your own data.** Almost every argument about this tradeoff is conducted on vibes and on published results from other people's targets. It does not have to be. **Instrument the program before arguing about it.** For every engagement, record four counts and one duration: how many strings were optimised locally; how many produced the behaviour on a held-out local model; how many produced it at a hosted endpoint in scope; how many of those were still working when re-tested; and, from a scheduled re-test cadence, how long each working string survived. After three or four engagements you have an empirical **transfer share** and an empirical **survival time**, and the allocation stops being a matter of taste. **Read those numbers against the reporting cycle.** If working strings survive less time than it takes to write and deliver the report, a transferred exhibit is stale on arrival and certainly cannot serve as the regression test someone re-runs each quarter. That does not make the work worthless; it means the *string* is the wrong unit of delivery. The durable unit is the behaviour class plus the recipe — objective, surrogate choice, ensemble composition, search settings — which lets someone re-derive a candidate against whatever build is live for the price of GPU hours. **Where the numbers mislead, and this is the part that decides budgets.** - *A denominator swap in the transfer share.* Teams quote transfer rates computed only over strings that already passed held-out validation. That figure measures the last stage of the funnel, not the return on the pipeline, and it can look excellent while the end-to-end yield per GPU-week is dismal. Always report the share over strings the search produced, and keep the funnel visible. - *Right-censoring in the survival time.* If you re-test quarterly you cannot observe a three-week half-life; every string that died in week four is recorded as "still working at last check" or as an unexplained gap. A cadence coarser than the decay you are trying to measure guarantees an optimistic answer. - *Spare capacity read as free.* Idle GPUs make the marginal run look costless. The real cost is the people: engineers who can define an objective, pick a surrogate and interpret a failed search are scarce, must be staffed continuously to stay sharp, and are the same people who could be doing query-driven work on the deployed systems that the report is actually about. **What each method uniquely buys.** Local search runs on hardware you own: no per-query cost, no rate limits, no unusual traffic against a vendor's monitoring, and it can run before the engagement window opens. Its output is *candidates*, plus class-level evidence about models you can never instrument. Query-driven discovery is the only thing that measures the **deployed** system — model plus system prompt plus guards — which is normally the thing with an owner who can act. The sane default is a portfolio: local search as a candidate generator, query work as the measurement instrument, split tuned by your measured transfer share rather than by whichever paper was published most recently. **The cut criteria, stated so they can actually fire.** Retire the transfer capability when (a) the measured transfer share stays near zero across several engagements, meaning the surrogate guessing is not working against this target population; (b) the targets you face reject at a surrounding classifier so consistently that model-level artefacts never reach a model — more GPU hours cannot fix a request that is never delivered; or (c) the scope of the work is the deployed product, so a model-level finding has no owner and no remediation path. Even then, keep a small standing capability and the tooling warm: a target population with locally runnable models, or a customer who deploys open weights themselves, brings the whole method straight back, and rebuilding the skill from zero costs far more than maintaining it.
- What is the durable unit of a finding if the string itself expires within weeks?The behaviour class, evidenced by a dated exhibit, plus the recipe — objective, surrogate choice and search settings — that lets someone re-derive a candidate against the current build.
- Your measured transfer share is decent but survival time is under a month. What changes in how you run the program?Keep the local search as a candidate generator, but stop treating strings as deliverables: report classes, attach dated exhibits with raw counts, and schedule re-tests rather than re-quoting the old rate.
saying these in an interview costs you the question
- Allocating effort from published results rather than from the program's own measured transfer share.
- Promising transferred strings as durable regression tests.
- Committing GPU capacity with no plan to measure whether transfer ever pays off.
- Framing it as an all-or-nothing choice instead of a portfolio with a measured split.
- Ignoring that the scarce resource is often the people who can run and interpret the search, not the GPUs.