Two near-duplicate engagement predictors flip between +40 and -38 across bootstrap refits — how do you fix the model?
answer
- Ask what decision the model serves first
- The sum is stable, the split is not
- Drop, combine, or report the joint effect
- Relabel the survivor as the shared effect
- Only new data truly separates them
basics
~20 sFirst decide what the model is for. If it only forecasts, the swings are harmless. If someone will act on a coefficient, the data cannot separate the two effects, so drop one, combine them into a single measure, or report the well-estimated joint effect instead.
solid answer
~50 sWild sign flips across resamples are the fingerprint of two predictors carrying one signal, and the first move is to establish what the model is used for. For pure forecasting, nothing is broken: the *sum* of the two coefficients is stable across refits even as the split swings, and predictions depend on the sum. If the coefficient itself feeds a decision, I would choose among three honest options. Drop one predictor on domain grounds and label the survivor as the combined effect rather than that feature's own effect. Or replace both with one deliberately constructed measure — a sum, an average, or a single well-defined engagement metric — so the model estimates one interpretable parameter. Or keep both and report the joint effect with its interval, which is genuinely tight. What I would not do is ship whichever refit produced signs that matched the story, or pretend a `+40` from an unidentifiable pair is a finding.
code
python · 24 linesimport random, statistics
def fit(rows):
x1 = [r[0] for r in rows]; x2 = [r[1] for r in rows]; y = [r[2] for r in rows]
a = [v - statistics.mean(x1) for v in x1]
b = [v - statistics.mean(x2) for v in x2]
c = [v - statistics.mean(y) for v in y]
s11 = sum(i * i for i in a); s22 = sum(i * i for i in b)
s12 = sum(i * j for i, j in zip(a, b))
s1y = sum(i * j for i, j in zip(a, c)); s2y = sum(i * j for i, j in zip(b, c))
det = s11 * s22 - s12 * s12
return (s22 * s1y - s12 * s2y) / det, (s11 * s2y - s12 * s1y) / det
random.seed(0)
data = []
for _ in range(200):
u = random.gauss(0, 1)
v = u + random.gauss(0, 0.01) # near-duplicate of u
data.append((u, v, 3 * u + random.gauss(0, 1)))
for _ in range(5):
sample = [random.choice(data) for _ in range(200)]
b1, b2 = fit(sample)
print(f"b1={b1:8.2f} b2={b2:8.2f} b1+b2={b1 + b2:6.2f}")go deeper
Recognise that violently swinging coefficients across refits point to two predictors carrying the same information, and that the total effect can be solid even when neither individual number is.
Be able to explain why the sum of the two coefficients is stable while each part swings, and to show it with a resampling check rather than asserting it from the standard errors.
Demonstrate that you pick a remedy from the model's purpose: leave a forecaster alone, but for an interpreted coefficient choose deliberately between dropping, combining and reporting the joint effect, and relabel what the survivor means.
Own the standard, not just this model. Define which outputs are decision-grade, require stability checks before any coefficient is quoted externally, and be willing to tell a stakeholder that the question cannot be answered with the data available.
## Reading the symptom Coefficients that leap from strongly positive to strongly negative across bootstrap refits, while the model's fit quality stays flat, is a near-conclusive signature: two predictors are measuring almost the same underlying thing. The fit can pay one and charge the other, in any of infinitely many nearly-equivalent ways, and each resample picks a different point along that trade-off. The crucial detail is what *does* stay stable. Because the two estimates are strongly negatively correlated, their **sum** is estimated precisely even though each part is not. The model as a predictor is untroubled; the model as an explanation is unusable. ## Step one: what is this model for? Everything downstream follows from this, and a lead who skips it is solving the wrong problem. - **Forecasting only** — the swings are cosmetic. Fitted values depend on the combined contribution, which is stable. Do not mutilate a working forecaster to make a coefficient table look calmer. - **Attribution, pricing, policy or a headline effect size** — the swings are fatal, because the number someone will act on is exactly the number the data cannot determine. - **Feature-importance narratives shown to stakeholders** — also fatal, and more insidious, because a plausible-looking ranking will be read as causal by an audience that never sees the interval. ## The honest remedies **1. Drop one predictor.** Simplest and often right. Choose on substantive grounds — which measure is better defined, more reliably logged, more comparable across segments — never by which one has the smaller p-value in this sample. The essential discipline is relabelling: the survivor's coefficient now represents the *shared* effect of both, and reporting it as that one feature's own effect is a misstatement. Fit quality typically barely changes, which is itself the proof that little information was lost. **2. Combine them into one measure.** Sum them, average them, or define a single composite engagement metric and use it alone. The model then estimates one parameter that is well identified, and you have made an explicit, documented modelling decision instead of letting the fit make an arbitrary one. The cost is that the composite must be defensible to the people who will read it — a sum of two quantities on different scales usually is not. **3. Keep both and report the joint effect.** Present the confidence interval for the combined effect, which is tight, and state plainly that the split between the two cannot be resolved. This is the most honest option and often the least popular, because it declines to answer the question that was asked. **4. Get data that separates them.** The only remedy that adds real information. An experiment, or a period or segment where the two metrics genuinely diverge, breaks the redundancy at the source. This is slow, and worth proposing when the attribution question is important enough to fund. ## What not to do - **Do not choose the refit whose signs match the expected story.** This is the failure mode that turns a diagnosable statistical limitation into a fabricated finding, and it is the specific thing an interviewer at this level is probing for. - **Do not iteratively delete predictors until a diagnostic table looks clean.** That optimises the diagnostic rather than the analysis, and each deletion changes every remaining number. - **Do not present a `+40` and a `-38` in a stakeholder deck with a caveat in the footnote.** Nobody reads the footnote; the numbers become the takeaway. - **Do not resolve it by adding more near-duplicates of the same behaviour** in the hope that averaging emerges by itself. It does not; the redundancy only deepens. ## Making the call durable At lead level the deliverable is not one fixed model, it is a rule the team applies next time: - Run a stability check — refit on resamples and record coefficient ranges — before any coefficient is quoted externally, not after someone questions it. - Establish that reported effects come with intervals, so an unidentifiable quantity announces itself instead of being rounded into a claim. - Keep a curated feature definition layer so two teams do not independently ship two names for one behaviour, which is where most of these pairs originate. - Write down which model outputs are *decision-grade* (interpretable, stable, defensible) and which are *forecast-grade* (accurate but not interpretable). Most of the damage from collinearity comes from a forecast-grade output being read as decision-grade. ## The summary a strong candidate gives Diagnose the redundancy, show that the sum is stable while the parts are not, ask what decision the model serves, then choose deliberately between dropping, combining, and reporting the joint effect — and be explicit that no analysis technique can extract a separation the data does not contain.
- If you drop one of the two predictors, how should the survivor's coefficient be described?As the combined effect of the behaviour both predictors measure, not as that single feature's own effect. The remaining predictor is standing in for both, and its estimate absorbs the shared contribution. Reporting it under the dropped feature's narrower label overstates what was estimated and invites decisions the data does not support.
- How do you decide between dropping one predictor and summing the two into a composite?By whether a sum is substantively meaningful. If the two are on the same scale and measure the same behaviour, a composite estimates one interpretable parameter and uses both signals. If the scales or definitions differ, the sum has no clean interpretation and dropping the weaker-quality measure, chosen on domain grounds rather than on p-values, is cleaner.
- A stakeholder wants the individual effect anyway. What do you tell them?That the data cannot answer it, and why: the two metrics almost never move independently, so no fit can attribute the outcome between them. Offer the combined effect with its interval, which is precise and decision-usable, and propose the experiment or the segment analysis that would actually separate the two if the distinction is worth funding.
- Why is a bootstrap refit a better stability check than looking at the coefficient's standard error alone?The standard error is a single summary that rests on the model's assumptions; a resample shows the actual spread of estimates you would have reported under slightly different data, including sign reversals. Seeing a coefficient move from strongly positive to strongly negative is far more persuasive to non-statistical stakeholders than a wide interval.
saying these in an interview costs you the question
- Picks the refit whose coefficient signs match expectations
- Deletes predictors until the diagnostic table looks clean
- Reports the unstable coefficient with a footnote caveat
- Never asks whether the model explains or predicts
- Keeps the survivor's original label after dropping its twin
- Claims a better fitting procedure could recover the split