Ranking candidate fraud signals by mutual information with the label, how would you decide which to keep?
answer
- the ranking is an input, not a policy
- pairwise scores double-count shared bits
- high cardinality inflates the estimate
- shuffled label as the control
- availability and ownership at scoring time
basics
~20 sTreat the ranking as one input to a decision, not the decision. Pairwise scores price shared bits only: they are blind to redundancy between candidates, inflated for high-cardinality columns, and silent about leakage, availability and the cost of obtaining a signal.
solid answer
~50 sA pairwise ranking answers one narrow question per candidate — how many bits does this column share with the label, alone, on this sample. Four things it cannot see usually decide the outcome. **Redundancy:** two near-duplicate columns both score high and the ranking credits their overlapping bits twice, so select greedily against the label *conditioned* on what is already chosen. **Estimation bias:** a column with almost one distinct value per row scores near the label's entropy on the sample and generalises nothing; the control is to re-score against a shuffled label, where a real signal collapses and an artefact does not. **Leakage:** the top scorer is often written after the outcome, so check when each value first exists. **Cost:** availability at scoring time, latency, who owns the upstream field, and whether the dependence survives the next quarter.
go deeper
Take away the caution rather than the procedure: a ranking of candidate signals is a starting list of hypotheses, and the highest number on it is not automatically the best thing to use.
Explain why pairwise scoring double-counts overlapping information between two similar candidates, and what a conditional score — bits added given what is already selected — does about it.
Demonstrate the controls you would actually run: a shuffled-label baseline for cardinality artefacts, a provenance check on the top scorers, and a time-ordered holdout with fields frozen to their decision-moment values.
Own the trade-off nobody measures: every kept signal is a standing dependency on an upstream producer, with monitoring and explanation costs attached, and its measured value decays as behaviour shifts. Set the cadence for re-measuring and the rule for retiring one.
## What the ranking actually prices Each row of the ranking is one number: the bits a single candidate column shares with the outcome label, computed alone, over one sample, at one binning. That is a genuine and useful measurement. It is also narrow enough that a top-k cut off the raw list is rarely the right policy. | The score prices | The score is silent about | |---|---| | Shared bits between one column and the label | Whether two candidates carry the *same* bits | | Dependence of any shape, not just linear | Which side is cause and which is effect | | The sample, at the binning chosen | Whether the same holds on data not yet seen | | — | Whether the value exists at scoring time | | — | What obtaining the column costs, in latency or in ownership | ## Four blind spots, and the control for each 1. **Redundancy between candidates.** Two columns that are near-copies of each other both score high, and a pairwise list credits the shared bits to each in full. The fix is to score **conditionally**: at each step, choose the candidate that adds the most bits about the label *given* the candidates already selected. A near-duplicate then scores near zero once its twin is in, and drops out by itself. The cost is real — each conditional estimate is computed on thinner slices of the sample, so the estimates get noisier as the selected set grows — which is why the loop is usually run to a modest depth rather than exhaustively. 2. **Cardinality inflating the estimate.** Estimated from counts, a column with nearly one distinct value per row leaves almost no uncertainty about the label within each cell, so the sample estimate climbs toward the label's entropy while the population value may be nothing at all. Anything identifier-like — a key, a hash, a timestamp at full precision — behaves this way. The control is to re-score against a **randomly shuffled label**: a genuine dependence collapses toward zero, an artefact of cardinality does not. 3. **Leakage.** The measure is symmetric, so a column derived from the outcome scores exactly as high as one that predicts it — usually higher, sitting near the label's entropy ceiling. The control is a provenance question, not a computation: at the moment a score must be produced, does this field already have a value, and is its writer upstream or downstream of whatever decides the label? 4. **Cost and durability.** None of these appear in the number at all. A column that scores well but is available only hours later, or is owned by a team that may redefine it, or exists for only a slice of the population, may be worth less than a lower-scoring one that is always there. ## A workable selection loop 1. Rank all candidates pairwise, and read the list as a set of **hypotheses**, not a decision. 2. Re-score everything against a shuffled label and discard any candidate whose score barely moves — that is the cardinality control. 3. Run the timeline check on everything near the top; anything written downstream of the outcome leaves the pool regardless of its score. 4. Select greedily with conditional scores, stopping when the marginal bits a further candidate adds fall into the noise of the estimate. 5. Re-check the surviving set on a **time-ordered holdout**, with each field frozen to the value it held at the decision moment, to confirm the dependence is not an artefact of the window sampled. ## The judgement a lead actually owns Beyond the mechanics there is a decision no measurement makes: how many signals the system should depend on at all. Every kept candidate is a permanent dependency on an upstream producer, something to monitor, and a piece of the system that has to be explained when a decision is challenged. A short list of stable, cheaply available, well-understood signals is often worth more in production than a longer list that scores better on a sample, because the marginal bits of the tenth candidate are usually small while its marginal operational cost is not. The second judgement is about drift. A dependence score is a photograph of one window. Adversarial behaviour in particular moves in response to whatever is being used, so a signal's measured value is not a constant of nature; the policy needs a re-measurement cadence and a rule for what happens when a signal's contribution decays. ## What an interviewer is listening for There is no single right answer here, and a candidate who produces a threshold ("keep everything above 0.1 bits") has missed the question. What earns the point is naming at least two blind spots with the control that exposes each, describing conditional rather than pairwise selection in a sentence, and treating availability, ownership and drift as first-class inputs alongside the measurement.
- Why does a near-unique identifier column score so highly, and what control exposes it?Estimated from counts, a column with almost as many distinct values as rows leaves essentially no uncertainty about the label within each cell, so the sample estimate approaches the label's entropy while the population value may be nothing. The control is to re-score against a randomly shuffled label: a genuine dependence collapses toward zero, an artefact of cardinality survives.
- What does conditioning the score on already-selected signals buy you?It prices the marginal bits a candidate adds rather than its bits in isolation, so a near-duplicate of something already chosen scores near zero and falls out on its own. The cost is that each conditional estimate is computed on thinner slices of the sample, so the numbers get noisier as the selected set grows and the loop has a practical depth limit.
- Is a higher-scoring signal always worth more than a lower-scoring one?No. The score says nothing about whether the value exists when a decision has to be made, how much of the population it covers, who owns the field upstream, or whether the dependence will hold next quarter. A slightly weaker signal that is always present and stable is frequently the better dependency to take on.
saying these in an interview costs you the question
- Takes the top-k of a pairwise ranking as the final set
- Ignores that two high scorers may be near-duplicates
- Trusts a near-unique identifier's high score as real signal
- Never checks whether a column exists at scoring time
- Assumes a dependence measured once holds indefinitely