A retrained ranker's per-segment quality report is all green — what does that establish about an adversary?
answer
- green in which segments, exactly?
- finer slices sit inside reported ones
- small segments are noisy both ways
- granularity trades against resolution
- absence of signal from a partial instrument
basics
~20 sOnly that no segment the report breaks out moved more than that segment's sample noise allows it to resolve. Damage confined below the reporting granularity, or inside a small noisy segment, reads green exactly as health does.
solid answer
~50 sA green per-segment report is a bound with two scopes attached, and both are usually left off when it is quoted. The first is **granularity**: it constrains only the segments actually broken out. An effect confined to a slice that sits inside a reported segment — a locale, a device class, one traffic source — is averaged in with everything else in that segment and does not move it. The second is **resolution**: each segment's number has sampling noise proportional to its size, so a small segment is green under mild damage just as readily as under health, and adding finer segments makes each one noisier. An adversary aiming at a narrow slice therefore does not need restraint to stay green; the report's structure does the hiding. What the green does establish is real but narrow: no broad, reported-segment-wide effect above noise. Report it in those words.
code
text · 11 linesoffline replay: candidate r-2481 vs current r-2477
segment n events quality delta alert
overall 12,400,000 0.741 -0.002 no
new users 910,000 0.712 -0.004 no
returning 11,490,000 0.743 -0.002 no
alert fires at delta <= -0.010 on any reported segment
not broken out: locale, device class, traffic source,
or any slice below 500,000 events
...go deeper
Know that a segment number is an average over everything inside that segment, so anything affecting only part of it barely moves the figure.
Be ready to explain the two scopes — which segments were printed, and how much noise each one's sample size carries — and why they limit the report in different ways.
Demonstrate the sign-off habit: convert a green table into a bound with its window, segments and alerting band named, and refuse to let it stand in as a statement about the whole model.
The judgment is where the reporting floor sits and what you do below it, since finer slicing runs out of traffic before it runs out of places to hide. Decide what you fund instead, and say what you are accepting.
## What is being read A per-segment quality report compares a newly retrained model against the previous one, broken down by a handful of segments a team decided to track. It is the most granular instrument most teams have, and it is the one that gets quoted as "we checked and everything is fine." Understanding what it can and cannot bound is the whole of this question. ## Scope one: granularity The report constrains movement in the segments it prints. Every slice finer than those segments is inside one of them, averaged with everything else there. If a segment covers eleven million events and the affected slice covers eighty thousand, an effect that is severe on that slice moves the segment's number by roughly the slice's share of it — a rounding error against the threshold anyone would alert on. This is why a narrow adversary does not need much restraint. They are not sneaking under a threshold by careful tuning; the report's own structure averages them away. Restraint is required only when the effect they want is broad. ## Scope two: resolution Each printed number is an estimate from a finite sample, so it carries noise proportional to that segment's size. A small segment's figure bounces around run to run, which means two things at once: it reads green under mild damage as easily as under health, and it also raises false alarms, so teams widen its alert band, which lowers its sensitivity further. Splitting into more, smaller segments does not give you a free increase in coverage — it trades granularity against resolution, and there is a floor set by how much traffic each slice actually gets. ## The direction of the inference Stated honestly, all green establishes: *no effect exceeding the resolvable noise of each reported segment, in the segments reported, over the compared window.* That is genuinely useful. It rules out a crude broad degradation. It does not rule out an effect that is confined, one that is inside an unreported slice, or one small relative to a noisy segment's band. The error to avoid is the inversion: reading the absence of a signal from an instrument as evidence about a thing the instrument does not measure. A green report is consistent with no adversary and with a competent narrow one, and by itself it does not discriminate between them at all. ## The conversation this leads to Asked what the report is worth next quarter, the useful answer is comparative rather than absolute. It will keep catching broad degradation and clumsy or unlucky campaigns, which is a real class. It will not start catching a confined effect, because that is a structural property of averaging and not a tuning problem. Buying more segments buys resolution down to the point where each slice's traffic runs out, and buys nothing below it. Anything below that floor needs a different activity — evaluating behaviour on inputs chosen because you suspect them, rather than on a sample of ordinary traffic — with its own scope and its own limits. ## Triage, when a slice does look off The symmetric error is over-reading a red. A single segment moving is far more often an ordinary cause — a traffic-mix change, an upstream logging change, a seasonal shift, a genuine data-quality bug — than an adversary. The useful discipline is to ask what would distinguish those, which usually means checking whether the movement tracks something exogenous, whether it persists across retrains, and whether the affected population has any relationship to who could write into the corpus. Concluding "poisoning" from a moved number alone is as bad as concluding "clean" from a flat one; both replace a bound with a verdict. ## What to say when signing off If you are the person whose name goes on the check, phrase the conclusion with its scope attached: which segments were compared, over what window, with what alerting band, and the explicit statement that effects confined below that granularity are not covered. That sentence costs nothing and prevents the report from being quoted later as an assurance nobody made.
- Why not simply add many more segments until nothing can hide?Because granularity trades against resolution. Each extra split shrinks the traffic behind every number, so its sampling noise grows, its alert band has to widen to avoid false alarms, and its sensitivity drops. There is a floor set by how much traffic a slice actually receives, and below it more segments buy nothing.
- One segment did move by more than the band. Is that a poisoning finding?Not on its own. Traffic-mix shifts, upstream logging changes, seasonality and ordinary data-quality bugs all produce this far more often than an adversary does. Check whether the movement tracks something exogenous, whether it persists across retrains, and whether the affected population overlaps with anyone able to write into the corpus.
- How would you word the sign-off so the report is not over-quoted later?Name the scope in the conclusion: these segments, this comparison window, this alerting band, no effect above the resolvable noise in each. Then state plainly that effects confined below that granularity are not covered, so nobody reads a green table as an assurance about the whole model.
saying these in an interview costs you the question
- All segments green means the model is clean
- More segments always means more coverage
- A small segment reading flat is strong evidence
- One moved segment proves an adversary
- The report covers slices it does not print