Procurement says the signed checkpoint's chain of custody is complete — how do you decide whether it ships?
answer
- verification finished; the question did not
- this is a trust decision now
- more paperwork answers a different question
- size it by consequence, not certainty
- spend at the deployment boundary
basics
~20 sTreat it as a trust decision, not a verification gap. Custody checks cannot answer a behavioural question, so decide by consequence: what a wrong output costs here, what you can constrain downstream, and what recourse you hold.
solid answer
~60 sThe first move is to say plainly what the custody evidence established — this is the vendor's file, published by the vendor — and that it makes no claim about behaviour. That reframes the meeting: nothing further is pending, so the residual is a trust relationship with the publisher, and demanding more artefact claims will not move it. Then decide on grounds a lead actually owns. Size the exposure by consequence: what does one wrong output in this deployment cost, who notices, and how quickly. Spend where it pays — constraining what an unreviewed model decision is allowed to trigger downstream buys far more than another ingest check, and a second independent signal on consequential actions is usually cheaper than any assurance about the weights. Convert what cannot be tested into allocated liability through the contract. Finally, record the decision as trust extended, with a named owner and a review point, so the organisation chooses it rather than inheriting it by default. Usually the answer is ship — knowingly.
go deeper
Recall that someone has to decide whether a borrowed model ships, and that a passing verification is an input to that decision rather than the decision itself.
Be able to explain to a non-specialist why a complete chain of custody leaves a behavioural question open, without overstating it into a reason to reject the vendor.
Show that you would size the exposure by what a wrong output costs in this specific deployment and constrain downstream consequences, rather than seeking assurance the artefact cannot provide.
Own the call and its record: name the residual, allocate it to somebody, use contractual recourse for what cannot be tested, and set the policy for which decisions may rest on inherited weights at all.
## The meeting you are actually in Somebody has presented a complete chain of custody for a checkpoint and is treating it as a security sign-off. The technical content of your answer is a single sentence — those checks establish which bytes and whose, and neither is a statement about behaviour — but the value you add as a lead is what you do with the room after saying it. The most common failure is not shipping something bad. It is letting an unanswered question be recorded as answered, so that nobody in the organisation ever knowingly decided anything. ## Step one: name the residual out loud State the position in words a non-specialist can carry: *the verification is finished and it succeeded; it tells us the file is the vendor's. Whether the model does anything the vendor did not disclose is a question those checks never addressed, and we cannot fully test it. We are choosing to rely on the vendor.* This is uncomfortable precisely because it is honest, and it is the sentence that changes the organisation's behaviour. An unnamed residual is accepted silently by everyone; a named one gets an owner, a review date, and a budget conversation. It also stops a predictable dead end. When a behavioural question is raised, the reflex is to demand more paperwork per artefact at ingest. More custody evidence answers custody questions better and answers this one not at all, so pushing that direction spends real effort and moves nothing. Say so early. ## Step two: size it by consequence, not by certainty Since you cannot reduce the uncertainty much, decide on the consequence side. - **What does a wrong output cost here?** A detector whose output queues a human review is a very different proposition from one that triggers an automatic action against a person. The same untested checkpoint is acceptable in the first and may not be in the second. - **Would anyone notice?** A failure mode nobody detects for a quarter is worse than a louder one. Ask what signal would surface an anomaly at all. - **How reversible is it?** If a bad period can be replayed, re-decided, or compensated, the exposure is bounded in a way that a permanent decision's is not. - **How exposed is the deployment to somebody who wants the condition to occur?** A model whose inputs are captured in a controlled environment differs from one whose inputs anyone can arrange to present. ## Step three: spend where the money moves the outcome The leverage is not on the artefact; it is around it. - **Constrain what an unreviewed model decision may trigger.** Caps on consequential actions, a human in the loop above a severity threshold, or a second independent signal for anything irreversible. These bound the damage regardless of why the model was wrong, which is the property you want when you cannot enumerate the causes. - **Instrument the deployment.** Per-slice monitoring in production continues the acceptance argument after ship, and it catches ordinary quality regressions too, which is how you fund it. - **Reduce dependence for the decisions that matter most.** You rarely have to choose between borrowed weights and none. Deciding which classes of decision may rest on an untested inherited model is a policy lever, and a cheaper one than assurance. ## Step four: use the levers that are not technical A contract can allocate the consequence of an undisclosed behaviour even though it cannot detect one. A vendor attestation that no undisclosed conditional exists is a liability instrument, not evidence, and it is worth obtaining for exactly that reason — provided you present it as such rather than filing it as a control. Exit options matter too: a deployment that can fall back to a previous model or a different supplier holds a different risk profile from one that cannot. ## Step five: write the decision down as what it is The record should say: custody verified; behaviour untested beyond the stated coverage; residual accepted by a named owner; compensating constraints listed; review at a stated point or on a stated trigger. That artefact is what makes the call auditable and revisitable, and it is the deliverable a principal is judged on here. ## And usually you ship A lead who blocks every third-party checkpoint has not managed the risk, they have exported the decision to whoever overrules them, and no organisation can fund building every model from scratch — which merely relocates the trust to your own data sources and pipeline anyway. The judgment being tested is not whether you can say no. It is whether the yes is a decision somebody made, sized, constrained and owns, rather than a verification result somebody misread.
- The vendor offers a signed statement that the model contains no undisclosed behaviour. What is that worth?It is a liability instrument, not evidence. It allocates consequence if the statement turns out false, which is genuinely useful and worth obtaining. It changes nothing about what you know, so present it in the risk record as recourse rather than filing it as a control that closes the finding.
- Does training your own model instead remove this exposure?It relocates it — to your data sources, your labelling supply chain and your own pipeline — and it costs money most programmes do not have. The realistic version of that lever is narrower: decide which classes of decision are allowed to rest on inherited weights, and fund independence only where the consequence justifies it.
- How do you keep this from becoming a blanket ban on third-party models?By tying the control to consequence rather than to provenance. Inherited weights behind a human review queue need a light touch; inherited weights driving an irreversible action against a person need constraints or a different sourcing decision. A single rule applied to both either blocks useful work or waves through the case that mattered.
saying these in an interview costs you the question
- Demands more artefact claims to answer a behavioural question
- Presents completed custody checks as a safety sign-off
- Blocks all third-party weights as a blanket policy
- Leaves the residual unnamed so it is accepted by default
- Files a vendor assurance letter as a technical control