skip to content

After fine-tuning an inherited checkpoint, a disclosed backdoor's success rate fell from 96% to 11% - what does that establish?

level: seniorimportance: should knowfreq 36%

answer

  1. ask what the number is attached to
  2. one key, one recipe, one input set
  3. who chooses the input at inference
  4. which row is doing the work

basics

~20 s

A measured reduction for one disclosed key under one adaptation recipe, on the inputs tested - not removal. It bounds no other key, and eleven percent against an adversary who chooses the input and retries is not small.

solid answer

~50 s

Read what the number is attached to. It is one key, disclosed by someone else, tested on one adapted copy produced by one recipe with a specific step count and pruning step, measured as a success rate over inputs you assembled. That is a real, useful result and it is not a claim of removal. Three things it does not establish: that any other key is affected, since a pathway you were never told about was never tested; that eleven percent is operationally safe, because the adversary chooses the input and can send many; and that the reduction is a property of fine-tuning, since in a table like this most of the drop usually arrives with the pruning row and that row carries its own accuracy cost. What you can report is a measured rate, a recipe and a scope - never a negative.

code

text · 7 lines
text
adaptation              params updated   steps   key tested   ASR     clean acc
published checkpoint    -                -       key A        96.4%   91.2%
frozen body + head      0.4%             2k      key A        93.8%   90.7%
full fine-tune          100%             2k      key A        41.5%   91.0%
full fine-tune + prune  100%             2k      key A        11.3%   90.1%
...
(no second key tested; key A disclosed by a third-party report)

go deeper

for a junior

Recall that a success rate is measured against a specific key on specific inputs, so it describes that test and not the model as a whole.

for a middle

Explain why a reduction is not a bound, and why the row that matches your production pipeline is the only one you may quote.

for a senior

Demonstrate the operating judgment: interrogate the scope columns, notice which stage produced the drop, and price a residual rate by what one activation is worth.

for a principal

Be ready to set the standard for how such results are written up, so a scoped empirical rate never leaves the building as a claim of removal.

## The result is real; the claim people make from it is not A measured drop from 96% to 11% is genuine evidence and worth having. The failure mode is the sentence that gets written underneath it: *the fine-tune removed the backdoor*. Everything useful in this question is in the gap between the measurement and that sentence. ## What the number is actually attached to Four scopes, all narrow: 1. **One key.** Somebody disclosed a key for this checkpoint family, which is the only reason you could measure anything. A conditional is fired by a feature the adversary chose; a key nobody published was never in your test set. Testing what you were told about tells you about what you were told about. 2. **One recipe.** A step count, a corpus, a learning rate, a freeze policy, and in this table a pruning stage. Change any of them and the residual rate changes. Another team running 'the same' full fine-tune on a different corpus can land somewhere else entirely. 3. **One input population.** A success rate is measured over inputs you assembled and keyed. A different population - longer inputs, a different domain slice, a slightly varied key - gives a different rate, and the adversary is not obliged to use yours. 4. **One direction.** The measurement establishes that this behaviour got harder under this treatment. It bounds nothing. There is no guarantee here, only an empirical rate. ## Why 11% is not the comfortable number it looks like Robust-accuracy intuition from average-case performance does not transfer, because this is not an average case. The adversary chooses the input, chooses when to send it, and can send many. A conditional that fires roughly one time in nine, at their discretion, on an input they construct, may be entirely sufficient for the payoff - especially where the payoff is a single retrievable document surfacing in an internal search index, or a single decision going their way. Ask what one successful firing is worth before deciding whether a rate is small. ## Which row is doing the work In a table shaped like the one above, the frozen-body row barely moves (93.8%), the full fine-tune roughly halves it, and the large drop arrives only with pruning. That matters for two reasons. First, if your production pipeline does not prune, the row you should be reading is 41.5%, not 11.3% - people quote the best row in the table and ship a different recipe. Second, the pruning row has its own bill: clean accuracy slipped, and on a per-slice basis a compression step usually costs rare classes and small subgroups more than it costs the average. You are trading measured task quality for an unquantified reduction in an unquantified risk, which is a decision, not a control. ## The reporting columns you would ask for If you were handed this table in a review, the questions are mechanical: which keys were tested and who disclosed them; how many inputs the rate was measured over; whether the deployed pipeline matches the row being quoted; whether clean accuracy is reported per slice as well as in aggregate; and whether anyone re-measured after the final production adaptation rather than a lab approximation of it. A row missing any of those is a number without a scope. ## The honest write-up Something close to: *under our production adaptation, the one publicly disclosed key for this checkpoint family fired on 11% of the keyed inputs we constructed, down from 96% on the published weights; this result covers that key only, at that recipe, and we have no test for undisclosed keys.* That sentence is defensible in front of anyone. 'The backdoor was removed' is not, and it is the sentence a reviewer is listening for.

  • Your production pipeline uses a frozen body, but the table's best row includes a full fine-tune and pruning - what do you report?
    The row that matches production, which is the 93.8% one, and you say so explicitly. Quoting a recipe you do not run is the most common way these tables mislead. If you want the lower number you have to actually adopt the recipe that produced it, and then absorb its accuracy cost, including the per-slice cost that the aggregate figure hides.
  • How would you decide whether an 11% success rate is acceptable?
    By the value of one firing, not by the rate. Ask what a single successful activation buys the adversary, how many attempts they can make undetected, and whether anything downstream would notice. If one firing surfaces a chosen document or flips one consequential decision, and attempts are unlimited and unlogged, 11% is effectively availability. If a firing is caught by a second signal, the rate matters more.
  • Would repeating the measurement with more keyed inputs strengthen the claim?
    It tightens the estimate of that rate, which is worth doing, but it does not widen the scope at all. The binding limitation is the single disclosed key, not the sample size. More inputs for one key gives you a more precise number about one key, and precision is often mistaken for coverage in write-ups like this.

saying these in an interview costs you the question

  • Reading a reduction as removal
  • Quoting the table's best row while shipping another recipe
  • Calling 11% safe without asking what one firing buys
  • Assuming an untested key behaves like the tested one
  • Ignoring the clean-accuracy cost of the pruning stage

context