You sign off a feature whose answers come from a generative step and will sometimes be wrong. What observation reverses that decision, and who watches?
answer
- One bad answer proves nothing here
- Compare against the figure you signed
- Nobody labels outputs in production
- A review level and a revert level
- A named person and a working fallback
basics
~20 sWrong answers are forecast, so one bad output reverses nothing. Write a rate on an observable signal — edits, discards, repeat attempts, escalations — with a review level, a revert level, a time window, a named person watching, and a fallback that already exists.
solid answer
~50 sA sign-off on a feature that is only usually right predicts wrong answers, so no individual wrong answer can falsify it. The reversal condition therefore has to be a **rate against the figure you signed off with**, measured on something observable in production, where nobody labels the outputs: the share of results people edit, discard, ask to be produced again, escalate to a person or report, plus how often the resulting action is undone downstream. Write two levels rather than one — a review level that opens an investigation and holds any further expansion, and a revert level that returns traffic to the previous path — with a window and a minimum volume so a handful of early reports cannot fire it and a slow rise cannot hide. Name the person who reads it, the fallback they invoke, and how long the heightened watch lasts.
code
pseudocode · 12 linessignal = share_of_results_edited_or_discarded
baseline = 0.11 # same signal, measured before release
window = last 24 hours
min_attempts = 500 # below this the rate is not read at all
every hour:
if attempts(window) < min_attempts: continue
observed = signal(window)
if observed >= baseline * 2.0:
revert() # route to previous path; named owner may act alone
else if observed >= baseline * 1.4:
open_investigation() # and hold the next exposure increasego deeper
Be ready to say why a feature expected to be wrong sometimes needs a rule agreed in advance about when it gets turned off, rather than a decision improvised when the first complaint arrives.
Explain what can be measured after release when nothing is labelled — edits, discards, repeat attempts, escalations, downstream reversals — and why each of them needs a baseline taken before launch.
Show that you would write two levels, a window and a minimum volume, name the person who reads the signal, and confirm the fallback exists and has been exercised before you sign anything.
Own the policy: what every released probabilistic feature must have in place to count as signed off, how long heightened watching lasts, and how you keep reverting an ordinary outcome rather than an admission of failure.
## Why a single wrong answer proves nothing The sign-off statement already predicts that the feature will be wrong sometimes. That has a consequence teams miss: after release, a report reading "the feature gave someone a wrong answer" is not new information and cannot reverse anything. If every reported error is treated as grounds to reconsider, the decision belongs to whoever complains loudest; if none is, the decision is never reconsidered at all. Both failures come from the same gap — nobody wrote down what *would* count as evidence that the forecast was wrong. The condition that closes the gap compares the world with the figure you signed off. You said roughly this share of attempts would return something wrong, under this input mix, with these consequences. The reversal condition names the observation that contradicts one of those clauses. ## What you can observe when nothing is labelled Production has no marking scheme. Nobody tells you which answers were right, so the condition has to be written on **proxies** — signals that move with wrong answers and are recorded anyway: - **Correction signals** — the share of outputs a person edits before using, discards, or asks to be produced again. - **Escape signals** — rapid repeat attempts, abandoning the flow, or falling back to the manual path the feature was meant to replace. - **Escalation signals** — contacting support, reporting the output, or a colleague reversing what was done. - **Downstream reversal** — how often the resulting record is amended, cancelled or refunded. - **Continuing judged sample** — a small, ongoing sample of real outputs judged by the same rule used before release. It is the only signal that measures the thing itself, which is why it earns its cost even at low volume. Every one of those has a rate you can compute before release from the attempts you already judged, and that gives the condition a baseline. Without a baseline the trigger is a number somebody invented. ## Writing the condition A usable condition has five parts: | Part | What it fixes | |---|---| | Signal | which observable the rate is computed on | | Two levels | a review level that opens an investigation, a revert level that pulls traffic back | | Window and minimum volume | over how long, and below what count the rate is not read at all | | Comparison | against the pre-release baseline for that signal, never against zero | | Action | what reverting concretely does, and who may do it without convening anyone | The two levels matter more than the exact numbers. A single threshold forces the watcher to choose between ignoring a worrying rise and withdrawing a feature on thin evidence. The review level buys the investigation; the revert level removes the argument. The window and the minimum volume prevent the two classic failures: firing on the first three reports from an unrepresentative early cohort, and never firing because a slow rise always reads as noise when it is checked daily. ## Who watches, and for how long Name a person, not a department. The condition should say who reads the signal, on what cadence, and who may invoke the fallback on their own authority. An owner recorded as a team has no owner. It also needs an end. Heightened watching is expensive and it decays, so agree a period: after it, either the feature moves into ordinary operations with the signal on a standing view, or the fact that it still needs somebody watching is itself a finding about whether it should have shipped in this shape. Where exposure is staged, the same condition governs each expansion. The rate is read at the current share of traffic before the next share is granted, and a review-level breach holds the expansion rather than merely being noted in passing. ## What makes reversal real A reversal condition is theatre unless three things are true, and they are worth checking while the recommendation is being written rather than during the incident: 1. **A fallback exists and has been exercised.** Reverting must mean routing to the previous behaviour, hiding the feature, or serving a plain message — not shipping a fix. If the only way back is a release, the condition cannot be met quickly enough to matter. 2. **The signal is instrumented before launch.** A proxy nobody records is a promise rather than a control, and instrumenting it afterwards means the baseline is already gone. 3. **Reverting is not a career event.** If withdrawal is treated as an admission of failure rather than as the plan working, the watcher will negotiate with the threshold instead of reading it. Saying in advance that reversal is an expected outcome at a stated rate is what keeps the control usable. Finally, separate two ways the forecast can fail: **the estimate was optimistic**, meaning the same inputs go worse than measured, and **the inputs moved**, meaning real traffic is not the mix you sampled. Both raise the same signal and need different responses, so the investigation the review level opens should establish which one it is before anything is changed.
- Nobody labels outputs in production. How do you tell whether the released rate matches what you signed off?Combine cheap proxies with a small continuing judged sample. Edits, discards, repeat attempts, escalations and downstream reversals are recorded anyway and can be baselined before release. Alongside them, judge a modest ongoing sample of real outputs by the same rule you used earlier — it is the only signal that measures correctness directly, and low volume is enough to catch a large move.
- Why write two levels instead of a single withdrawal threshold?Because one number forces a bad choice: ignore a worrying rise, or withdraw the feature on thin evidence. A review level opens an investigation and holds any further expansion while it runs; a revert level removes the judgement call entirely. The gap between them is where you establish whether the estimate was optimistic or the real input mix has changed.
A rule that reverses on any single wrong answer is a smoke alarm mounted over a stove: it will be disabled within a week, and then nothing is watching at all.
saying these in an interview costs you the question
- Treats any single wrong answer as grounds for withdrawal
- Writes the trigger as when users start complaining
- Names a department rather than a person as the watcher
- Sets a threshold with no pre-release baseline for the signal
- Has no way back except shipping a fix