In Rosenbaum bounds for a matched study, what does the sensitivity parameter Gamma represent?
answer
- a bound on treatment odds within a matched set
- Gamma of one means as-if random
- raise it until significance is lost
- worst case over hidden assignment patterns
- amplification splits it into two strengths
basics
~20 sGamma is the largest factor by which two units with identical measured covariates may differ in their odds of receiving treatment. Gamma equals 1 means assignment within matched pairs is effectively random; larger values allow more hidden bias.
solid answer
~50 sRosenbaum bounds ask how far a matched study can depart from random assignment before its conclusion breaks. Within a matched set, units look identical on everything you measured, so under the no-hidden-bias assumption each is equally likely to be treated. Gamma relaxes that: it caps the ratio of the odds of treatment between two matched units at Gamma, so Gamma of 1 is as-if random and Gamma of 2 means one unit could be twice as likely to be treated for reasons you never observed. For each Gamma you compute the worst-case p-value over all hidden assignment patterns consistent with that bound, then raise Gamma until that bound crosses your significance threshold. The reported result is a sentence like `the finding survives a 1.6-fold departure from random assignment` - larger tipping points mean more robust findings, and there is no fixed value that counts as safe.
go deeper
Recognise the shape of the result: a matched study reports the size of departure from random assignment its conclusion can withstand, with 1 meaning perfectly random within pairs.
Explain the mechanics - a bound on the ratio of treatment odds within a matched set, a worst-case p-value computed at each value, and the tipping point where significance is lost.
Show judgment about plausibility. Argue from what drives selection in the domain whether the reported tipping point is a real defence, and know that the basic bound is conservative about outcome strength.
Own design sensitivity. Choose comparisons and matching ratios that will tolerate more hidden bias before results are collected, rather than negotiating over a weak tipping point afterwards.
## What matching assumes and what Gamma relaxes Matched designs pair each treated unit with one or more untreated units that look the same on measured covariates. The hope is that within a pair, who got treated is effectively a coin flip. If that were literally true, the study would be a randomised experiment stratified by the matching variables, and the usual permutation reasoning would apply exactly. It is not literally true. Something unmeasured may still push one member of a pair toward treatment. Rosenbaum's sensitivity analysis parameterises exactly that departure. Write the odds that a unit receives treatment; for two units in the same matched set, the assumption is that the ratio of their treatment odds lies between `1/Gamma` and `Gamma`. - Gamma = 1: identical odds within every matched set. This is the no-hidden-bias case. - Gamma = 1.5: one matched unit may be up to 50 percent more likely to be treated than its partner for unobserved reasons. - Gamma = 3: a threefold unobserved imbalance is permitted. Gamma is a bound on the *design*, not a description of reality. Nobody knows the true value; that is the point. ## How the analysis runs Fix a Gamma. The bound admits a whole family of hidden assignment mechanisms, and each would produce a different p-value for the outcome comparison. Rosenbaum's method computes the extremes of that family - most usefully the largest p-value any mechanism within the bound could produce, the worst case for your claim. Sweeping Gamma upward, this worst-case p-value rises monotonically. The reported number is the Gamma at which it first exceeds the significance threshold. That number is the headline: `the result is insensitive to hidden bias up to Gamma of 1.6`. Below that departure from randomness, no assignment pattern consistent with the bound could have produced the observed data by hidden bias alone. Above it, some pattern could. ## Reading the number honestly Three readings are wrong and worth naming. First, Gamma is not an estimate of the bias present. The study does not measure Gamma; you choose values and see what happens. Reporting `Gamma is 1.6 in this study` as a finding misstates the method. Second, there is no universal threshold. Whether 1.6 is impressive depends on the domain. Where treatment selection is driven by rich measured data - administrative eligibility rules, say - a 1.6-fold unobserved imbalance may be genuinely implausible. Where selection is driven by clinical judgment or motivation, a Gamma of 1.6 is trivially attainable and the result is fragile. As with any sensitivity analysis, the number is only as good as the argument attached to it. Third, a low Gamma is not proof of confounding. It means the finding is fragile to it, which is different. ## Gamma versus confounder strength The basic Gamma bound is deliberately worst case with respect to how strongly the hidden variable relates to the outcome - it effectively allows the unobserved variable to be a perfect predictor. That makes a small Gamma sound worse than it is, because it takes both an unobserved imbalance in treatment *and* a strong relationship with the outcome to actually create bias. The amplification technique addresses this. A single Gamma is translated into a curve of equivalent pairs: a large unobserved imbalance in assignment combined with a weak outcome relationship, or a small imbalance combined with a strong outcome relationship, both producing the same worst-case sensitivity. Reporting the amplification makes the result far easier to argue about, because reviewers can point at a specific pair and say whether such a variable is plausible in this domain. ## Where it fits Rosenbaum bounds are the natural sensitivity analysis for matched-pair or matched-set designs analysed with permutation-style tests, and they sit alongside other tools rather than replacing them: threshold-style summaries answer the same question on the risk ratio scale for regression-adjusted estimates, and falsification checks catch a broken pipeline that no sensitivity parameter would flag. A study that reports a matched estimate, its Gamma tipping point with an amplification, and a pre-specified falsification check has said everything an observational design can honestly say. One design consideration matters more than the analysis. Rosenbaum's notion of *design sensitivity* is that some designs simply tolerate more hidden bias than others before their conclusions break - a stronger dose contrast, a more homogeneous comparison group, or multiple untreated matches per treated unit will typically raise the achievable Gamma. That means sensitivity is partly something you buy at design time, not merely something you report after the fact. If you know your result will be scrutinised for hidden bias, choosing the comparison that would survive a larger Gamma is a better investment than any post-hoc adjustment.
- A matched study loses significance at Gamma of 1.05. What does that tell you?The finding is extremely fragile: an unobserved factor making one matched unit merely 5 percent more likely to be treated could account for it. That is well within what any realistic unmeasured variable could do, so the estimate should not be treated as evidence of an effect without a stronger design.
- Why does amplification make a Gamma value easier to interpret?The raw bound is worst case in how strongly the hidden variable predicts the outcome, so it understates robustness. Amplification restates one Gamma as a set of equivalent pairs - imbalance in assignment against strength on the outcome - letting a reviewer judge whether any real variable occupies one of those combinations.
- Can you increase the Gamma a design can survive before collecting data?Yes, and this is the idea of design sensitivity. A sharper treatment contrast, a more homogeneous comparison pool, or extra matched controls per treated unit typically raise the tipping point. Sensitivity is partly a design choice made in advance, not only a statistic reported afterwards.
It is a stress rating on a bridge. You do not measure the load that will actually cross it; you report the heaviest load it is certified to carry, and let engineers judge whether real traffic stays under that.
saying these in an interview costs you the question
- Says Gamma is the amount of hidden bias the study contains
- Treats a fixed Gamma value as universally acceptable
- Confuses Gamma with the confounder's association with the outcome
- Claims a high Gamma proves the effect is causal
- Thinks a low Gamma demonstrates a confounder exists