A vendor's model checkpoint is yours to inspect byte for byte — why does that not reach behaviour they trained in?
answer
- the reviewable object is missing
- no named routines inside a tensor
- what would you diff it against?
- retraining is not bit-reproducible
- review turns into search over inputs
basics
~20 sA checkpoint is a large array of numbers with no named routines and no clean reference copy to compare against. A behaviour trained in by the publisher is spread across the same parameters as everything else, so inspection has nothing to localise.
solid answer
~50 sReviewing a source patch works because there is a readable change and a reference to diff it against. A checkpoint offers neither. The parameters encode the whole function jointly, so a conditional the publisher trained in does not sit in an identifiable region; it is expressed in the same numbers that produce the ordinary behaviour. There is also nothing to diff against: retraining the same architecture on the same data is not bit-reproducible, and any legitimate fine-tune moves essentially every parameter, so a comparison shows uniform difference and localises nothing. Model capacity is far larger than the task needs, which leaves representational room for a conditional that costs nothing measurable elsewhere. That leaves exactly one channel to behaviour — feed the model inputs and read the outputs — which converts a review problem into a search problem over an input space you do not control and did not choose.
go deeper
Recall that a checkpoint is numbers rather than readable instructions, and that you cannot see what a model does by looking at it — you have to run it.
Explain the two missing ingredients that make code review work: a legible artefact and a reference to diff against. Be able to say why retraining does not supply the second.
Be ready to shut down the retrain-and-diff proposal quickly in a design discussion and redirect the effort into evaluation, naming what evaluation can and cannot cover.
Own the consequence for sourcing strategy: assurance for borrowed weights comes from the publisher relationship and from testing, so decide what you are willing to depend on rather than commissioning inspection work that cannot pay off.
## What review of a code change actually depends on When a security reviewer approves a source change, two things are quietly doing the work: a **readable artefact** — statements a human can follow — and a **reference** to compare against, namely the code as it stood before. Provenance tells you the change came from the right party; the review reaches behaviour because the object of review is legible and the difference is small and localised. A published checkpoint has neither property, which is why the reviewing habit transfers so badly to borrowed weights. ## The artefact is not legible A checkpoint is a set of numeric arrays. There are no named routines, no branches to read, no comment explaining an odd constant. Behaviour is *distributed*: the mapping from input to output is produced jointly by the parameters, and there is no sub-region that means one thing while the rest means another. A conditional response the publisher fitted into the model — the model behaves normally except under some condition they chose — is carried by the same parameters that produce the ordinary behaviour. Nothing about a parameter's value marks it as belonging to that conditional. This is also why the intuitive fix of looking for anomalous parameters does not survive contact. Deep models have far more capacity than their task requires; there is representational room to spare, so a conditional need not distort the parameter statistics in any way an inspector would flag, and normal training already produces wide, irregular distributions of values. ## There is nothing to diff against The second instinct is comparative: retrain the model yourself, or obtain a clean copy, and diff. Both fail for the same structural reason. - **Retraining is not bit-reproducible.** Nondeterministic parallel reductions, data ordering, hardware and library differences all shift results. Two honest runs of the same recipe differ in essentially every parameter. - **Legitimate change is not sparse.** Any fine-tune the publisher performed — and publishers fine-tune constantly — moves nearly all parameters. A diff against any plausible reference is dense by construction. So the comparison you would want, of the form *these few numbers changed and here is why*, is not available at any price. The output of the diff is uniform difference, and uniform difference localises nothing. ## What that leaves One channel remains: run the model on inputs and observe outputs. This is a real channel — behaviour is genuinely observable — but notice what it changes about the task. Review is bounded work over an artefact you hold. Evaluation is *search* over an input space that is effectively unbounded, and the conditions that matter were chosen by the publisher, not by you. You are looking for a response that occurs on inputs you were never told to try, in a model that is by design accurate everywhere else. That asymmetry — cheap for the publisher to install, expensive for the recipient to find — is what makes inherited weights a distinct problem rather than an instance of ordinary artefact review. ## Consequences worth stating in an interview - **Byte-level access is not the constraint.** Teams sometimes ask for the raw file as though obtaining it settles matters. You may hold every byte and still be unable to say what the model does. - **Format does not decide this either.** A weight file that cannot execute anything on load is a good thing for entirely separate reasons, and it changes nothing about whether the trained function has a conditional in it. - **Openness narrows trust rather than removing it.** Publishing the corpus and the training recipe moves the question from the weights to the data and the run. Both are large, and connecting corpus to checkpoint requires reproducing the training, which is expensive and almost never done by the recipient. That is a real improvement in what could in principle be examined, not a conversion of trust into a check. - **Small models are a partial exception.** For a small model over a handful of features you can sometimes characterise the learned function directly. That exception does not scale, and interviewers like candidates who name it and then say why it does not help for anything with real capacity. The conclusion a strong candidate reaches out loud: with borrowed weights, assurance about behaviour is bought by evaluation and by the trust relationship with the publisher, and inspection of the artefact contributes almost nothing to it.
- Would a much smaller model make weight-level inspection workable?At the margin, yes. For a small model over a handful of features you can sometimes characterise the learned function directly and reason about its whole domain. It does not scale: anything with real capacity has representational room to spare for a conditional, and no parameter-level reading localises it.
- Does publishing the training corpus and recipe close the gap?It narrows it rather than closing it. The question moves from the weights to the data and the run, both of which are large, and linking a corpus to a specific checkpoint requires reproducing training — expensive, rarely done, and not bit-reproducible anyway. It gives you something examinable in principle, not a check you actually perform.
- Why do people expect a conditional to show up as an anomalous layer or neuron?Because they are importing an intuition from code, where a malicious behaviour is a distinct block somebody wrote. In a trained model the same parameters carry the ordinary and the conditional behaviour jointly, and normal training already produces wide, irregular parameter distributions, so there is no anomaly of the kind they are picturing.
saying these in an interview costs you the question
- Claims a planted behaviour shows up in a parameter diff
- Expects to localise the behaviour to one layer or neuron
- Thinks the file format is what blocks inspection
- Assumes retraining yields a comparable clean reference
- Treats obtaining the raw weights as settling the question