Your policy for third-party model checkpoints names only the file format, and teams read a pass as approval to ship. What do you change?
answer
- one rule, two questions
- write coverage as non-coverage
- no scan makes borrowed weights clean
- tier by consequence, not uniformly
- someone signs the residual by name
basics
~20 sSplit one rule into two questions. Write the format requirement as what it covers - code execution when the file is read - add a separately owned question about what the weights do, and name who accepts the residual you cannot test away.
solid answer
~40 sThe defect is not the rule, it is that one rule sits where two questions belong and a single pass reads as a verdict. I would rewrite the policy so each control states its scope *and* its non-coverage in its own text: the format requirement closes code execution when the file is read and makes no claim about behaviour. Then add the second question explicitly - what would make us willing to serve weights somebody else trained - and accept that it has no clean answer. The real levers are narrowing eligible sources for anything whose output has consequence, behavioural acceptance tests keyed to our own harm cases with their limits reported, and containment downstream so the model's output is not treated as authority. Someone senior signs the residual, by name, per consequence tier.
go deeper
Take away the shape of the problem: a control that closes one risk will be quoted later as though it closed all of them, so what a check covers has to be written down beside its result.
Be able to say why more scanning is the wrong answer here - a clean behavioural result bounds only what was searched, which is a weaker claim than the format rule already provides.
Show you would tier the bar by consequence and would fix where files are first opened, since evaluation hosts sit upstream of every production gate a policy imposes.
Own the residual explicitly: decide which sources are eligible, where containment absorbs what detection cannot, and who signs for weights you cannot establish anything about.
## When one rule sits where two questions belong This is a governance failure more than a technical one, and it is worth naming as such. The format rule is a good rule. It closes a real class of attack structurally rather than by detection, it is cheap to enforce, and it needs no judgment at the point of use. The problem is that it is the *only* rule, so it has become the thing people point at when asked whether a downloaded checkpoint is safe - and it does not answer that question. It answers a narrower one that happens to be the one somebody wrote down. ### Change one: controls state their non-coverage The cheapest and highest-leverage edit is textual. Every control in the policy gets its scope written as a pair - what it closes and what it says nothing about - and the review artefact prints per-channel results instead of one verdict. A record that reads *loader channel: closed; weight behaviour: not assessed* is read correctly by the next person; a record that reads *APPROVED* is not. The wording is doing real work here, because the person reading it six months later is not the person who ran the check, and they are usually deciding something bigger than the checker was. This also protects the rule itself. Right now, when somebody eventually demonstrates that a format-compliant checkpoint can still misbehave, the political outcome is often that the format rule loses credibility - even though it never claimed what it was used to claim. ### Change two: name the second question and admit it has no clean answer Add the behaviour question to the policy explicitly. Then resist the instinct to answer it with more scanning. A behavioural probe bounds what it searched - the input families and conditional shapes you looked for - and a clean result from it is a weaker claim than the format rule's, while sounding like a stronger one. Funding detection here buys the appearance of coverage. What is actually available: - **Source eligibility.** For deployments where the output has consequence, narrow which publishers are usable at all. This is a decision about who you are willing to depend on, not a test, and it should be described that way. - **Acceptance testing keyed to your own harm cases.** Probe the specific behaviours you would care about, and report the result with its bound attached. It is real evidence for a narrow claim. - **Containment downstream.** Constrain what the model's output is permitted to decide without a check. This is the only lever that holds regardless of what is in the weights, which makes it the best place for money. - **Explicit acceptance.** For the remainder, a named owner signs. ### Change three: tier by consequence, or the policy will be routed around A uniform bar fails in both directions - too heavy for the experiments that make up most usage, too light for the handful of deployments that matter. Tiering is what makes the policy survivable: for low-stakes internal work, the format rule plus a contained load *is* a reasonable bar, and the policy should say so rather than leaving teams to infer that they are non-compliant and stop reading. For high-consequence deployments, the second question must be answered with behavioural evidence or with containment, and the eligible source list shrinks. Most work lands in the first tier, so tiering buys speed rather than costing it. ### Change four: fix the scope of *where*, not just *what* A policy about production models misses the machine that actually opens the file. Checkpoints are downloaded and read during evaluation, on laptops and build runners that often hold broader credentials than the serving host. The policy has to say which hosts may open unreviewed files at all. This is one of the few places where a small operational change removes a large amount of exposure. ### What you say to the teams Be direct about what has changed and what has not. The format rule still stands and nothing they do today becomes non-compliant. What has changed is that a pass no longer reads as clearance, because it never established what it was being used to establish - and for the small set of deployments where it matters, there is a second question with an owner and a bar. Framing it as *the same rule, now with its coverage written down* gets far less resistance than framing it as a new gate. ### The judgment you are being scored on An interviewer is listening for three things: that you refuse to answer a coverage gap with more detection; that you locate the fix partly in wording and artefacts rather than entirely in tooling; and that you say plainly who holds the risk that remains. Anyone can propose a scanner. The principal-level answer is to be honest about what cannot be established and to make somebody own it.
- Teams will say this slows them down. What do you actually let through unchanged?Almost everything. For internal and low-stakes work, the format rule plus a contained load stays the whole bar, and the policy says so explicitly. Only deployments whose output drives money, safety or a customer-visible decision have to answer the behaviour question, with evidence or with containment. Most work sits in the first tier, so tiering buys speed rather than costing it.
- Why not fund stronger behavioural scanning and keep a single rule?Because a clean behavioural result bounds only what was searched - the input families and conditional shapes you looked for - and never the model. It buys a weaker claim than the format rule while sounding like a stronger one. The same money spent on containing what the model's output can decide holds regardless of what is in the weights.
- How do you word things so a pass stops reading as approval?State each control's scope and its non-coverage in the same sentence, and stop the review artefact printing a single overall verdict. A record saying "loader channel closed, behaviour not assessed" is read correctly; one saying "approved" is not. The wording carries the load because the reader later is rarely the person who ran the check.
- Who ends up owning the part you cannot test away?A named person per consequence tier, not the reviewer and not the policy. The reviewer can only report what was checked; the policy can only say what evidence is required. Somebody with authority over the deployment has to accept that weights trained by an outside party are being served on incomplete evidence, and re-accept it when the tier changes.
saying these in an interview costs you the question
- Answers a coverage gap with more scanning
- Bans third-party checkpoints instead of tiering
- Leaves a single overall pass/fail verdict in place
- Assigns no owner to the untestable residual
- Scopes the policy to production and ignores build hosts