A partner's records already skewed your live forecaster and the fix lands at the next retrain — who decides whether it keeps serving?
answer
- It is an acceptance decision, not a fix
- Price the alternative before recommending it
- Some parts are not up for discussion
- Refuse the all-or-nothing framing
- Name the date the model becomes clean
basics
~10 sNot the responder. The owner accountable for the downstream service accepts it, on a written statement of what is skewed, how wide, for how long, and what the alternative costs.
solid answer
~50 sThis is an acceptance decision, not a technical one, and it belongs to whoever owns the courier-staffing outcome the forecasts feed. What you owe them is precise: which slice is affected, how far off it is, the date the deployed model becomes clean, and what serving the alternative costs — because a global fallback to a naive baseline may well be worse for everyone than a scoped skew is for one region. Two things must not be traded away. First, the write channel gets suspended regardless of the serving decision, or the next snapshot is contaminated again. Second, the record must say that no request-path control removed this, so nobody later reads a deployed gateway rule as a closure. Decide up front what evidence would show the retrain worked — a check on the affected slice against a pre-contamination reference, not flat aggregate error.
go deeper
Understand that removing bad training data and having a clean deployed model are two different milestones, and that someone senior has to accept the gap between them.
Be able to lay out what the accountable owner needs: which slice is affected, by how much, until when, and what the fallback would cost by comparison.
Show that you would close the write channel first regardless of the serving decision, and that you would refuse to record a request-path control as the remediation for a train-time root cause.
Own the acceptance framing: supply priced options rather than a binary, name the date the deployed model becomes clean, fix the verification evidence before the retrain, and make sure the residual exposure cannot be read later as closed.
## Why this is not the responder's call The technical facts here are settled by the time the question arrives: an adversary stood before training, the effect is in the deployed weights, a request-path control cannot reach it, and the earliest clean model is one training cycle away. What remains is a choice between two harms, and both land on someone else's outcome. Keeping the forecaster serving means known-wrong numbers for a known slice for a known window. Pulling it means every consumer runs on a fallback that is worse for everyone. Choosing between those is an accountable owner's decision — the person answerable for the staffing outcome — and the responder's job is to make it decidable rather than to make it. ## What the owner needs in front of them A useful acceptance package is short and has no adjectives in it: - **Scope.** Which regions, product families or time windows are affected, and how you determined that. If you cannot define the slice, say so, because it changes the decision entirely — an unbounded skew and a localised one are not the same risk. - **Magnitude.** How far the forecasts move on that slice against a pre-contamination reference. - **Duration.** The date the deployed model becomes clean, tied to a specific run, not "after the next retrain". - **The alternative's cost.** What a fallback baseline does to forecast quality, globally or on the affected slice, in the same units. - **What is not remediated.** Explicitly: the contaminated rows are out of the dataset and the write channel is suspended; the serving weights still carry the effect. ## The parts that are not negotiable Two items should not be presented as options. **The write channel closes now.** Whether or not the model keeps serving, leaving the partner's records flowing into the next snapshot means the audited data goes stale and the retrain relearns the effect. This is where the decision becomes genuinely commercial — the partner may be a large supplier, and suspending an integration has a relationship cost that the responder does not get to spend. Escalate it as a separate decision with its own owner, but do not let it be deferred, because everything else depends on it. **The finding is not closed on a request-path control.** A gateway rule may be deployed as a stopgap and it may even help against something else, but recording it as the remediation for this root cause encodes exactly the misconception that caused the argument. Write the residual exposure down with an expiry date attached. ## Scoping beats an all-or-nothing choice The binary framing — keep serving or pull it — is usually the wrong one, and the principal move is to refuse it. If the skew is confined to a slice, route that slice to a baseline and leave the rest of the estate on the model. If the consumer can be bounded instead, cap how far a staffing decision may move from the previous week without a human look. Both are compensating controls, both are cheaper than a global fallback, and both need to be recorded as temporary with the same expiry as the residual exposure. Offering the owner three options with costs is a materially better conversation than offering two. ## Answering "is it fixed?" honestly Someone will ask, and the answer has three clauses that must all appear: the contaminated rows are removed from the dataset; the write channel that supplied them is suspended; the deployed model still reflects them until the run on a named date. Collapsing that into "yes, remediated" is the failure this whole leaf exists to prevent, and it is a failure of wording rather than of engineering — which is why it is so easy to commit under pressure. ## Decide the verification before the retrain, not after Commit in advance to what would show the retrain worked. Aggregate error holding flat is not it: a localised effect barely moves an aggregate, so flat aggregate error is what you would observe whether or not the effect survived. The check is on the affected slice against a pre-contamination reference. Agreeing that before the run removes the temptation to accept the first flat dashboard as proof, and it forces the scope question to be answered early, when it is still cheap. ## What this question is really testing Whether you can hold two things at once: that the technical conclusion is not in doubt, and that the decision arising from it is not yours. A candidate who spends the answer re-litigating the mechanism has missed it. A candidate who says "we took it offline" without asking what the fallback costs the consumer has swapped one unpriced harm for another. The expected shape is a decision made by the accountable owner, on numbers the responder supplied, with the non-negotiable parts already done and the residual exposure written down where it cannot be misread later.
- The partner is your largest data supplier. How does that change the call?It changes who signs it, not the facts. Suspending the write channel has a commercial cost the responder cannot spend, so it escalates as its own decision with its own owner. What must not happen is quiet deferral: if the channel stays open, the audited snapshot is stale and the retrain relearns the effect, so the deferral has to be recorded as an accepted risk with a date.
- What would you accept as evidence the retrain actually removed the effect?A check on the affected slice against a reference from before the contamination. Aggregate error holding flat proves the effect was localised or that the adversary preserved overall quality — not that it is gone. Agree the slice and the reference before the run, because defining them afterwards invites picking whichever cut looks clean.
- Legal asks whether the issue is fixed. What is the honest answer?Three clauses, all of them: the contaminated records are removed from the training data, the channel that supplied them is suspended, and the model in production still reflects them until the training run on a named date. Anything shorter is a claim the evidence does not support, and it is the claim people most want to make.
saying these in an interview costs you the question
- Takes the serving decision as the responder alone
- Pulls the model without pricing the fallback
- Presents keep-serving or pull-it as the only options
- Defers suspending the write channel to keep a partner happy
- Reports the finding remediated once the rows are deleted