Risk asks whether fine-tuning your inherited encoder removed anything planted in it - what do you commit to?
answer
- you cannot evidence a negative here
- state scope, not a verdict
- change the question to blast radius
- somebody must own the residual
basics
~20 sCommit to what you tested and what it covers, never to removal. Absence of a conditional keyed to a feature you do not hold is not demonstrable, so the decision to own is what the model's output may reach.
solid answer
~50 sThe honest statement is scoped: here is the publisher we inherited from, the adaptation we ran, the keys we were able to test and the rates we measured - and we have no test for a key nobody has disclosed. Refusing to say 'removed' is the whole job, because a conditional fires on a feature the adversary chose and your corpus and evaluation set both contain no instance of it. Then move the decision where it can actually be made: what the encoder's output is permitted to decide unaided, whether any consequential action rests on it alone, which publishers you are willing to inherit from at all, and whether the consequence justifies re-pretraining from a corpus you control. Further fine-tuning is the cheapest-looking option and the weakest; a named owner accepting a scoped residual in writing is worth more than another training run.
go deeper
Understand that nobody can test for a key that has never been disclosed, so 'we found nothing' is not the same as 'there is nothing'.
Be able to write the scoped statement: publisher, adaptation run, keys tested, rates measured, and what was not covered.
Show that you would move the discussion from certification to blast radius, and that you can rank the available responses by evidence produced per unit cost.
Own the call: choose which bill the organisation pays - product capability, engineering velocity or compute - and make sure a named person accepts a written residual rather than a reassuring sentence.
## Why the question as asked cannot be answered Risk wants a yes or no on 'did adaptation remove anything planted'. The reason no competent answer is a yes or no: a planted conditional fires on a feature chosen by whoever trained the published weights. Your fine-tuning corpus contains no instance of it, your evaluation set contains no instance of it, and any test you can construct covers only keys somebody has already disclosed. Absence of evidence here is manufactured by the same gap that let the conditional survive in the first place. And the adversary's own position is symmetric: publishing a checkpoint is a bet, not a guarantee. They did not control your corpus, steps, freeze policy or pruning. Neither of you knows. The difference is that only one of you is being asked to sign something. ## What you can say, and should say in exactly these terms - **Provenance of the artefact and the adaptation actually run** - which publisher, which adaptation style, what share of parameters was updated, how many steps, whether pruning was applied. - **What was tested** - which disclosed keys, over how many inputs, with the measured rates before and after. - **What was not tested, stated plainly** - no test exists for an undisclosed key, and a clean evaluation on ordinary inputs is not evidence about a conditional. - **The residual, named** - a stated probability-free acknowledgement that an inherited conditional may be present and that our adaptation applies pressure to it that we did not measure and cannot bound. That write-up is defensible in front of anyone, including someone who can compel an answer later. 'We fine-tuned it on clean data so it is gone' is the sentence that becomes a problem, because it was never supportable and everyone technical in the room knows it. ## Then change the question The decision a lead actually owns is not certification, it is exposure. Four levers, roughly in order of cost-effectiveness: **Bound what the output can reach.** An encoder powering an internal document-search index is far less dangerous than the same encoder gating an access decision or an automated action. Ask what a single anomalous output could cause with no second signal and no human in the path, and remove the paths where the answer is unacceptable. This is usually the cheapest control and the only one that works against a key you cannot enumerate. **Decide the inheritance policy.** Which publishers the organisation is willing to build on, and whether a widely adopted, long-lived, independently re-derived artefact is treated differently from a checkpoint that appeared last month. This is a policy call with a real cost - it narrows what your teams may use - and it belongs to a lead, not to a reviewer. **Monitor for the payoff rather than the key.** You cannot alert on a key you do not hold, but you can often notice its consequence: a document surfacing in results it should not rank for, a decision distribution shifting for a narrow slice of inputs. Weak, but it is the only detection that does not require knowing the trigger. **Re-derive the weights.** Pretraining from a corpus you control removes the inheritance entirely, and it costs what it costs. This is justified by consequence, not by anxiety: it belongs where the model gates something you would not be willing to explain having outsourced. ## The option to argue down 'Just fine-tune it harder' will be proposed, because it is cheap, familiar and feels like action. Say clearly what it is: unmeasured pressure from an objective aimed at something else, which does reduce disclosed planted behaviours and does not bound anything. Spending a second training budget on it converts money into a slightly better feeling and no new evidence. If the organisation wants evidence, the money goes to testing disclosed keys and to bounding blast radius, in that order. ## Who absorbs what Every honest option here has a bill and it lands on someone: narrowing the encoder's authority costs product capability, an inheritance policy costs engineering velocity, re-pretraining costs compute and calendar. Deciding which of those the organisation pays - and getting a named owner to accept the remaining residual with the scope written down - is the principal-level act. The technical finding is that you cannot certify; the leadership act is choosing what you do instead and making sure the decision has an owner rather than a comfortable sentence.
- Someone proposes another, longer fine-tune as the mitigation. What do you say?That it is pressure, not a control. It does measurably lower disclosed planted behaviours, so it is not worthless, but it produces no evidence and no bound, and it costs a real training budget. If we are spending money for assurance, testing every disclosed key and bounding what the encoder's output may decide unaided both buy more per pound than a second training run does.
- When is re-pretraining from a corpus you control actually the right call?When the consequence of one anomalous output is something you would not be willing to explain having outsourced to an anonymous publisher - a gating decision, a safety-relevant classification, a regulated determination. For an internal search ranker it is almost never worth the compute and calendar. Scale the response to what a single firing could cause, not to how unsettling the threat sounds.
- How do you describe the residual so it survives an audit two years later?Write down the publisher and artefact identity, the adaptation actually run in production, the keys tested with their sources and measured rates, the explicit statement that undisclosed keys are untestable, the compensating limits on what the output may decide, and the named owner who accepted it, with a date. Anything vaguer becomes an argument about what people believed at the time.
saying these in an interview costs you the question
- Signing off that clean fine-tuning removed the risk
- Treating an untestable question as an answerable one
- Funding another training run instead of bounding exposure
- Leaving the residual unowned and undocumented
- Scaling the response to fear rather than to consequence