skip to content

How would you set an organisation's default significance level when error costs differ sharply by decision?

level: principalimportance: nice to knowfreq 31%

answer

  1. a default is a coordination device
  2. argue classes, not individual cases
  3. decide before anyone sees the number
  4. cost, recoverability, and who pays
  5. audit realised outcomes, not intentions

basics

~20 s

Keep one published default for consistency, then allow documented exceptions by decision class, each justified in advance by the cost and recoverability of a false alarm versus a miss. The policy's real job is preventing thresholds chosen after the result is known.

solid answer

~50 s

I would publish a single organisational default so that most analyses are comparable and nobody negotiates a threshold per project, then define a small number of **decision classes** that may depart from it: irreversible high-cost commitments get a stricter level, cheap reversible screens that feed a confirmatory stage get a looser one. Every departure carries a written justification recorded before data collection: what a false alarm triggers, what a miss triggers, whether either is caught downstream, and who absorbs each cost. The governance point matters more than the number. A default without pre-registration just moves threshold-shopping into the exception process, so I would tie the level to the analysis plan, review exceptions at the class level rather than case by case, and periodically audit realised outcomes — how often flagged results failed to replicate, how often things the screens cleared came back as problems.

go deeper

for a junior

You are not expected to set policy. Know that the significance level is chosen rather than derived, and that whoever chose it should have written down why before the data arrived.

for a middle

Be able to explain what a change in the organisational default would do to both error rates across many analyses, and why the choice cannot come from the mathematics alone.

for a senior

Show that you would tie the level to the analysis plan and refuse a departure requested after the result is known. Bring evidence about which errors have actually hurt your team.

for a principal

Own the trade-off publicly: define a small set of decision classes, defend the risk appetite each encodes to non-statisticians, and build the audit loop that tells you when the classes are wrong.

## Why a default exists at all A significance level encodes how much false-alarm risk an organisation will tolerate in exchange for how much miss risk. In principle each decision deserves its own answer. In practice, letting every analyst derive one per project has two failure modes: results stop being comparable across teams, and the threshold quietly becomes negotiable after the fact. A published default solves both cheaply — it is a coordination device before it is a statistical statement. ## Why a single default cannot be the whole policy The costs genuinely differ. Consider two decisions in the same company. One: a screen that flags candidate issues for a second, more careful analysis. A false alarm here costs one extra investigation and is corrected immediately; a miss is never revisited. Two: a decision to withdraw a shipped product. A false alarm here is enormously expensive and public; a miss usually means waiting for more evidence. Forcing both to the same threshold means at least one of them is holding the wrong risk. So the policy needs a default plus a mechanism for principled departure. ## Decision classes, not case-by-case negotiation The design I would argue for is a small, closed set of decision classes — perhaps three or four — each with its own level and its own written rationale: - **Screening / triage**, feeding a stricter confirmatory stage: looser than default, because misses are unrecoverable and false alarms are absorbed downstream. - **Standard analysis** informing a reversible decision: the default. - **Irreversible or externally visible commitment**: stricter than default, because acting on a signal that is not real is the dominant cost. Classes beat per-project negotiation for a reason that is organisational rather than statistical: a class is argued once, in the abstract, before anyone knows which side of the line their own result will land on. Case-by-case exceptions are argued by the person who already saw the number. ## The threshold is a policy statement about who absorbs error The cleanest way to explain this to non-statisticians is a fixed-hardware screening system. At an airport scanner with the detection technology held constant, every threshold setting that catches more concealed weapons also stops more harmless passengers, and every setting that reduces passenger friction also lets more real threats through. No setting is "correct". The choice states which cost the institution will bear and on whose behalf, and it is legitimate for that answer to differ between a domestic terminal and a high-threat route. Leadership's job is to own the position, not to pretend the trade-off can be engineered away by tuning the threshold — the only way to improve both rates is better hardware, which in the analytics analogue means more data, better measurement, or better design. ## Governance is the part that actually binds A number in a wiki changes nothing on its own. Three mechanisms make it real: 1. **Pre-specification.** The level and its class live in the analysis plan, timestamped before data collection. This is the entire defence against choosing the threshold to match the result, which silently inflates the real false-alarm rate far above the stated one. 2. **Class-level review.** Departures are approved as classes by a standing group, not defended per analysis by the analyst who wants one. 3. **Outcome audit.** Track what happened: how often flagged findings failed to hold up on re-analysis, how often things a screen cleared surfaced later as real problems. If the first number is high the classes are too loose; if the second is high they are too strict. Without this loop the policy is untested opinion. ## What to say about the number itself Resist the temptation to make the answer be "use 0.05" or "use 0.01". The interviewer is testing whether you understand that no threshold is derivable from statistics alone — it requires a loss function the organisation has to supply. If pushed, the defensible line is: keep the conventional default where costs are roughly symmetric and reversible, because convention has real coordination value, and spend your influence on the exception classes and the pre-specification discipline, where the actual risk lives. ## Failure modes to name - **Silent drift.** Teams gradually loosen thresholds because loose thresholds produce more publishable-looking wins. An audit catches this; a policy document alone does not. - **Strictness as theatre.** Mandating a strict level everywhere looks rigorous and hides its cost, because the effects it causes you to miss are invisible in every report you read. - **Threshold as a substitute for design.** Any conversation that ends in "let's adjust alpha" when the honest answer is "this study is too small" has traded a design problem for a reporting one.

  • Why approve exceptions as decision classes rather than case by case?
    Because a class is argued in the abstract, before anyone knows which side of the threshold their own result falls on. A case-by-case exception is requested by the person who already saw the number, which turns the exception process into threshold-shopping with a review meeting attached.
  • How would you tell whether the thresholds a policy sets are actually well calibrated?
    Audit outcomes rather than intentions. Track how often flagged findings failed to hold up when re-examined, and how often issues a screen cleared resurfaced later as real problems. A high first rate says the levels are too loose; a high second rate says they are too strict. Without that loop the policy is untested opinion.
  • How do you explain a deliberately loose threshold to an executive who reads it as lower standards?
    Frame it as which mistake the organisation has chosen to bear. A screen that must not let real problems through accepts more false alarms on purpose, because each one is resolved by the next investigation while a miss is never revisited. Strictness is not free rigour; it buys quiet at the price of undetected problems.

saying these in an interview costs you the question

  • Answers with a single number and no policy
  • Lets each analyst choose a level per project
  • Approves exceptions after the result is known
  • Claims a stricter threshold is always more responsible
  • Never audits whether the chosen levels worked

context