Two moderation layers in the same stack each block about 90% of a 500-case attack corpus when measured in isolation. Why can you not conclude that dropping either one costs roughly ten points of coverage, and what measurement answers the question properly?
answer
- marginals do not determine the joint
- unique catch, not catch rate
- same-lineage layers correlate
- the neither cell is the real finding
- zero unique catch is not proof of uselessness
basics
~20 sBecause the two 90s may cover the same 450 cases. What matters is per-case overlap, not per-layer totals. Record every layer's verdict for every case in one table, then count the cases only one layer caught. That unique-catch column is the layer's marginal contribution; the aggregate rate hides it.
solid answer
~60 sTwo per-layer rates are marginal totals over the same cases, and marginals do not determine the joint. If the layers agree case for case, the stack blocks 450 and dropping either costs nothing — the survivors are the same 50 either way. If they are perfectly complementary, the stack blocks all 500 and dropping either costs 50. Same two 90s, opposite decisions. The measurement is the per-case matrix from the isolation runs: rows are corpus cases, columns are layers, and you count the four cells — caught by both, by A only, by B only, by neither. **A layer's unique catch is the only number that estimates what removing it costs.** Two cautions before that number ships. Layers built the same way — same style of classifier, same training data lineage — tend to be highly correlated, so redundancy is the default expectation, not the surprise. And the unique-catch count is measured on *your* corpus; it estimates the cost of removal against attacks like the ones you sourced, and says nothing about a class you did not include.
go deeper
Should at least see that the same cases might be caught by both, so 90 and 90 do not imply the stack catches 100 or that ten points are at stake.
Names the per-case matrix and unique catch as the right measurement, and can describe the two extreme joints that fit the same marginals.
Adds why correlation is the default for same-lineage layers, that the residual cell is the finding, and that unique catch is corpus-relative.
Turns it into a removal decision: unique catch against latency, spend and false positives, plus the redundancy value that survives a bad update or a provider outage.
### The error, stated precisely Two per-layer catch rates are **marginal totals** over the same set of cases, and marginals do not determine the joint distribution. "Layer A blocks 450 of 500" and "layer B blocks 450 of 500" are two facts about two columns; the question "what happens if I delete a column" is a fact about the rows they share, which neither total contains. The two extremes fit the same numbers. If the layers agree case for case, they block the same 450 and miss the same 50: the stack catches 450, and dropping either one costs **zero** coverage. If they are maximally complementary, A's 50 misses are exactly the cases B catches and vice versa: the stack catches all 500, and dropping either costs **50** cases, ten points. Identical marginals, opposite removal decisions. Anything in between is possible, so the honest statement from two 90s alone is "the stack blocks between 450 and 500", which is not a basis for touching anything. ### The measurement that answers it The per-case matrix from the isolation runs — one boolean per (case, layer), never a per-layer aggregate. For a pair, collapse it into a contingency table: ``` B blocks B misses A blocks a b <- b = A's unique catch A misses c d <- d = the stack's residual ``` Read it as: `a + b + c` is stack coverage. `b`, the cases **A alone** caught, is what you lose by deleting A; `c` is what you lose by deleting B; `a` costs nothing either way, because B still covers those cases once A is gone. `d` — the cases nothing stopped — is the most valuable cell in the engagement and the only one that is a *finding* rather than accounting. With more than two layers, compute it per subset, because the quantity that matters is always "unique to this layer given the others that remain". ### What drives the overlap, and what it costs to learn Predict before you measure: layers that share a detection principle share blind spots. Two classifiers trained on overlapping public safety data, or an input and an output screen that both key on the same surface vocabulary, land near the all-agree corner, and redundancy is the default expectation rather than the surprise. Genuine complementarity comes from layers that differ in kind — one reading intent in the request, one reading the realised content of the reply. The matrix either supports the defence-in-depth argument for this specific stack or refutes it; nothing else does. The cost is that this number cannot be bought cheaply. You need a per-layer verdict on every case, which means the isolation runs: baseline plus one pass per layer, each paying generation and classifier calls across both the attack and the benign half. There is no way to compute the joint from end-to-end runs, because short-circuiting deletes exactly the `a` column — the cases the first layer caught are precisely the cases the second was never asked about. ### Where the numbers mislead **Small counts.** Unique catch is usually a handful of cases. On 500 cases, `b = 7` against `c = 4` is not "A contributes nearly twice B" — the difference is inside sampling noise, and ranking layers on it is reading a coin. Report the counts with an interval, and refuse to order layers whose intervals overlap. **Corpus relativity.** Unique catch estimates what removal would have cost *against the attacks you sourced*. A layer with zero unique catch is not proven useless; it is unexercised in the direction where it would have mattered. Say this on the same line as the number, not in a footnote. **Redundancy that the matrix cannot see.** A layer with zero unique catch still covers the case where its neighbour ships a bad model update, is misconfigured during a deploy, or its hosted provider has an outage. Availability and change-risk are not in a block-rate table, and a report that recommends deleting a layer on unique catch alone has handed the owner half the picture. **False positives.** A layer's cost is not only latency and spend; it is the benign traffic it stops. The benign half belongs in the same matrix, and the layer worth arguing about is the one with small unique catch and large benign impact. ### What you would check That the matrix was built from isolation runs and not reconstructed from end-to-end logs; that `not reached` is a distinct value never folded into `not blocked`; that no isolated run rewrote the text a later layer read; that each unique-catch count carries its interval and its corpus caveat; and that the residual cell — the cases nothing stopped — leads the report, because it is the only column that changes what the team builds next.
- One layer shows zero unique catch across your corpus. Do you recommend removing it?Not on that evidence alone. State that it caught nothing your other layer missed on this corpus, then weigh its cost against what it covers when the other layer is down, misconfigured or newly updated. Removal is the owner's call with that framing.
- Which cell of the two-layer table do you lead the report with?The one where neither layer blocked. Everything else is attribution accounting; that cell is the actual gap, and it is the only column that changes what the team builds next.
Two nets that each catch 90% of the fish tell you nothing about what the pair catches: if the mesh is the same size they catch the same fish, and the second net is insurance rather than coverage.
saying these in an interview costs you the question
- Adding or subtracting layer catch rates as if they were independent.
- Recommending removal of a layer purely on a low aggregate rate.
- Computing overlap from end-to-end runs, where the short-circuited cells are missing.
- Ignoring that a redundant layer still protects against the other layer's misconfiguration or outage.
- Presenting unique-catch numbers without saying they are relative to this corpus.