skip to content

How does the file-drawer problem bias a published literature, and can you detect it?

level: seniorimportance: nice to knowfreq 26%

answer

  1. nulls never leave the drawer
  2. selection on the outcome
  3. surviving effects are too large
  4. effect size against precision, expect symmetry
  5. registered reports accept before results

basics

~20 s

Null results stay unpublished in the file drawer, so the visible literature is selected on significance: too many positives, and effects too large. A funnel plot of effect size against precision exposes it: missing small null studies appear as asymmetry.

solid answer

~50 s

The file-drawer problem is selection on the outcome: studies that reach significance get written up and accepted, studies that do not are quietly shelved. The published record is then a biased sample of the work actually done, and it is biased in a predictable direction — too many positives, and effect sizes systematically larger than the truth, because a small study only clears the significance bar when noise happens to help. The standard diagnostic is a funnel plot: plot each study's effect estimate against its precision. Under no bias you expect a symmetric inverted funnel, with imprecise small studies scattered widely and symmetrically around the true value. Publication bias hollows out one corner — the small studies with null or negative estimates are missing — producing visible asymmetry. Asymmetry is suggestive, not proof, since genuine heterogeneity between studies can also produce it. The structural fixes are registered reports and preregistration, which commit a journal or a team to the question rather than the answer.

go deeper

for a junior

Know the basic shape: results that find nothing are less likely to be published, so what you read is a filtered sample and the effects in it look stronger than they really are.

for a middle

Be able to explain both consequences, especially the magnitude one: a small study reaches significance only on a favourable draw, so surviving estimates sit in the tail. Describe what a funnel plot puts on each axis and why symmetry is expected.

for a senior

Show you read literature as evidence with a selection process attached. Discount small striking findings, weight large pre-specified studies, and name the caveats on funnel asymmetry rather than treating it as a verdict.

for a principal

Own the incentive design. Argue for registered-report-style commitments, internal results registries and a culture that publishes well-powered nulls, and be ready to say what that discipline costs in speed and in headlines.

## What the file drawer is Rosenthal's file-drawer metaphor: for every study that appears in print, some number of similar studies sit unwritten in a drawer because they found nothing. If publication depends on the result, then the published literature is not a sample of the studies conducted — it is a sample **selected on the outcome**, which is the classic recipe for a biased estimate. The selection has two distinct consequences, and strong candidates name both. **Too many positives.** If a large number of teams independently test null hypotheses, a fraction will hit significance by chance alone. When only the hits get published, the literature is populated by exactly those chance hits, each looking like a legitimate discovery. **Effects that are too large.** This one matters more and is missed more often. A small, noisy study reaches significance only when the sampling variation happens to push its estimate away from zero. So conditioning on significance conditions on a favourable draw, and the surviving estimate sits in the tail rather than at the truth. A literature assembled from surviving estimates therefore reports effects meaningfully larger than reality, even for effects that are genuinely non-zero. ## Funnel plots The standard visual diagnostic across a set of studies on one question. For each study, plot its effect estimate on one axis against a measure of its precision on the other — larger samples or smaller standard errors sit at the top, small imprecise studies at the bottom. With no bias, the picture is a symmetric inverted funnel. Precise studies cluster tightly around the common effect at the top; imprecise studies scatter widely, but symmetrically, to both sides. That symmetry is the signature of a complete record. Publication bias eats one lower corner. Small studies that found a null or opposite-signed effect never entered the record, so the bottom-left of the funnel is sparse while the bottom-right is full. The plot looks lopsided, and the average effect across studies is dragged upward by the surviving small studies. **Read funnel asymmetry carefully.** Asymmetry is evidence, not proof. Other explanations exist: real heterogeneity where smaller studies genuinely differ (different populations, more intensive interventions), quality differences correlated with size, or simply too few studies for the shape to mean anything. Treat asymmetry as a prompt to ask what is missing, not as a verdict. ## The replication evidence The Reproducibility Project in psychology took roughly a hundred published findings and repeated them with high-powered direct replications. Only around a third to forty percent produced a statistically significant result in the same direction, and the replication effect sizes averaged roughly half the originals. Comparable exercises in other empirical fields have landed in similar territory. What that number does and does not mean matters. It does **not** mean two-thirds of the studies were fraudulent, or that any single failed replication proves the original wrong — replications have their own sampling variation and can differ in population or procedure. What it does show is a literature whose reported effects are systematically larger than what careful repetition finds, which is precisely the fingerprint of selection on significance combined with flexible analysis. ## Why it compounds with analytic flexibility Selective publication and selective analysis multiply. An analyst with many defensible analysis choices can usually find one that reaches significance; the journal then selects among those already-selected results. Each filter passes through the estimates that noise flattered most. The reader sees the end of a long selection chain and has no way to reconstruct its length from the paper. ## Fixes that actually work - **Preregistration.** A time-stamped, public analysis plan filed before the data are seen. It fixes the hypothesis and the analysis, and it leaves a record even if the paper never appears. - **Registered reports.** The stronger version, and the one that attacks the file drawer directly: the journal reviews the introduction and method *before* data collection and grants in-principle acceptance. Publication then depends on the question and the design, not the result — so a null cannot be filed away. - **Study registries and results databases.** Making the existence of a study public at launch means a missing write-up is visible as a gap rather than invisible. - **Publishing nulls, and valuing them.** A well-powered null is genuinely informative about the world, and treating it as a non-result is what fills the drawer in the first place. - **Reading the literature as a selected sample.** Practically: assume published effects are upper bounds, weight large pre-registered studies far above small surprising ones, and be sceptical of a striking effect that exists in one small study. ## In an interview Expect this in research-adjacent or data-science roles as "why should I not just trust the published effect size?" The two-part answer — the record is selected on significance, so both the count of positives and the magnitude of surviving effects are inflated — plus one detection method and one structural fix is a complete answer.

  • Besides publication bias, what else can make a funnel plot asymmetric?
    Genuine heterogeneity is the main alternative: small studies may differ systematically from large ones in population, intensity of the intervention, or measurement quality, producing real differences in effect that mimic missing studies. Chance also matters when only a handful of studies exist. Treat asymmetry as a prompt to ask what is missing rather than as proof of suppression.
  • If only about a third of studies replicate, does that mean two-thirds were wrong?
    No. A failed replication is itself one study with sampling variation, and it may differ in population or procedure from the original. The defensible reading is aggregate: a literature whose replication effects average roughly half the originals is one where reported magnitudes are systematically inflated. That points at selection on significance and analytic flexibility, not at widespread fraud.
  • Why is a registered report a stronger fix than preregistration alone?
    Preregistration fixes the analysis but does not guarantee the study is published, so a null can still end up in the drawer. A registered report has the journal review the question and method before data collection and grant in-principle acceptance, so the paper appears whichever way the result falls. That removes the selection step itself rather than just documenting the analysis.

saying these in an interview costs you the question

  • Thinks publication bias only inflates counts, not effect magnitudes
  • Reads funnel asymmetry as conclusive proof of suppression
  • Says a failed replication proves the original study was wrong
  • Assumes meta-analysis automatically corrects for missing studies
  • Treats a well-powered null result as an uninformative non-result

context