skip to content

A team accepts third-party checkpoints only in a tensors-only weight format - what does that remove?

level: middleimportance: must knowfreq 55%

answer

  1. one channel closed, one untouched
  2. the container changed, not the contents
  3. the numbers cross over unchanged
  4. answers the loader question only
  5. flat metrics are the attacker's design goal

basics

~20 s

It removes code execution when the file is read, and nothing else. Weights are unchanged by the format holding them, so a network trained to misbehave is exactly as backdoored once it is stored as plain tensors.

solid answer

~40 s

The rule closes one channel completely and the other not at all. Requiring a format that cannot execute when it is read means opening a downloaded checkpoint is no longer an execution event, which genuinely kills the class of attack where merely reading a published file gets code onto your host. It does nothing to the numbers. A network whose weights carry a conditional behaviour the publisher trained in is exactly as backdoored after conversion, because a format change rewrites how tensors are laid out, not what they are. So a review that checks the format and stops has answered *can this file run code on us* and has not begun on *does this model do what we think*. Those two questions need different evidence - one about the loader, one about behaviour.

code

text · 7 lines
text
checkpoint-review: community-hub / text-to-image-v2
  weight format          tensors only, no executable objects    PASS
  contained load         no process spawn, no network egress    PASS
  eval set (1k prompts)  quality within 1% of baseline          PASS
  behaviour on chosen inputs   -- not assessed --
  ...
  verdict                APPROVED FOR PRODUCTION

go deeper

for a junior

Remember the direction: a file format is the container and the weights are the contents. Converting a checkpoint changes where the numbers live, never what they do.

for a middle

Be able to say what the rule removes structurally - the load is no longer an execution event - and then name, unprompted, the question it has not started on.

for a senior

Spot the supporting checks that are secretly the same channel: a contained load and a flat aggregate metric both leave the weights question exactly where it was.

for a principal

Insist that a review record report per-channel coverage rather than one verdict, because a single pass will be read downstream as clearance for risks nobody examined.

## What a format rule buys, exactly This is the single most common substitution error around inherited weights, and it is worth stating in the sharpest possible form: **a weight format is a property of the container, and a backdoor is a property of the contents.** Changing the container leaves the contents alone. That is not a subtlety, it is arithmetic - a conversion writes the same numbers into a different layout, and a model is its numbers. ### The half that is genuinely closed Give the rule full credit for what it does. If reading a file cannot construct arbitrary objects, then an adversary who publishes a checkpoint no longer gets anything from you merely by being downloaded. That was a real and cheap attack: the requirement was only that some process opened the file, which happens on laptops and build runners long before any review. Removing that requirement removes the class - not detects it, not makes it unlikely, removes it. Detection-shaped controls that scan a file for dangerous constructs are strictly weaker, because they bound the constructs the scanner knows to look for. A format that has no execution semantics is a structural fix and should be described as one. ### The half that is untouched What the rule does not do is say anything about behaviour. The adversary here is the *publisher* - they trained the weights, so they could have trained a conditional into them: normal output everywhere you will look, publisher-chosen output on inputs the publisher controls. Nothing about that behaviour lives in the file format. It lives in the values, and the values survive any conversion you perform on the way in. Worse, the two most common supporting checks a review adds are usually the same channel again in disguise: - **Loading the file in a sandbox and watching for activity** probes the loader channel a second time. A backdoored network spawns no process, opens no socket, and writes no file when it is read. It is inert until asked for a prediction. - **Running the evaluation set and seeing quality within a percent of baseline** does not probe the weights channel either, because preserving ordinary quality is the attack's design goal. A model that got visibly worse would be rejected. Flat aggregate metrics prove the publisher preserved them; they bound nothing about inputs your evaluation set does not contain. So a review can accumulate three green results and still have answered exactly one question. ### How to state the coverage honestly The useful discipline is to write each control's scope as its **non-coverage**, in the same sentence: | The control | What it closes | What it says nothing about | | --- | --- | --- | | Tensors-only weight format required | Code execution when the file is read | Anything the weights do | | Contained load, no observable activity | Observable loader behaviour, under those conditions | Anything the weights do | | Aggregate quality within tolerance | A visible degradation attack | Behaviour on inputs the publisher chose | | Behavioural probing on chosen inputs | The input families and trigger shapes you searched | Whether the file executes on load | The last row is the mirror image and matters just as much: a behavioural probe that came back clean tells you nothing about the loader. Teams that invest heavily in model-behaviour review sometimes make the opposite substitution and accept a file format they never examined. ### What would move the second needle Honestly, less than people want. You cannot re-derive somebody else's weights, so the available evidence is behavioural: probing the model with inputs shaped like the harm you actually care about, and looking for conditional structure in how it responds. Every such result is bounded by what you searched. That is a weaker claim than the format rule provides, and the right move is to report it as weaker rather than to let a clean result inherit the format rule's certainty. The reason this matters commercially is that the residual is real and someone has to hold it. A review that produces one overall verdict hides that; a review that reports per-channel results makes it visible, and visible residual risk is the kind somebody can decide about. ### The one-line version *The format question and the behaviour question are different questions. Closing the first is cheap, structural and complete. The second has no equivalent answer, and pretending otherwise is what the format rule keeps getting used for.*

  • The same team also loads the file under containment and sees no activity. Does that add anything about the weights?
    No - it probes the same channel again. A contained load observes what happens when the file is read, which is the loader question a second time. A backdoored network spawns nothing and calls nothing on load; it is inert until someone asks it for a prediction on an input the publisher chose.
  • Their evaluation shows quality within one percent of baseline. Why is that not evidence about the weights?
    Because preserving average quality is a design goal of the attack, not luck. A conditional behaviour is only useful if the model looks ordinary everywhere else, so flat aggregates are what a competent publisher engineered. They prove the publisher preserved them, and bound nothing about inputs your evaluation set does not contain.
  • You cannot re-derive the weights. What evidence would actually move your confidence about behaviour?
    Only evidence that touches behaviour: probing the model against inputs shaped like the harm you care about and looking for conditional structure in its responses. Every such result bounds what you searched - the input families and trigger shapes you looked for - and nothing outside them. That is a weaker claim than the format rule buys, and should be reported as weaker.

saying these in an interview costs you the question

  • Calls a safe-format checkpoint a clean model
  • Treats format conversion as sanitising the weights
  • Cites flat aggregate metrics as evidence against a backdoor
  • Assumes a contained load covers model behaviour
  • Reads one control's pass as a whole-file verdict

context