skip to content

Most metric-dip investigations on your team find nothing — how do you set the bar for investigating?

level: principalimportance: nice to knowfreq 33%

answer

  1. Investigation time is a finite budget
  2. Cost of a miss versus a false alarm
  3. Not every metric deserves an alarm
  4. Agree the rule before the dip
  5. Track what alarms actually turned out to be

basics

~20 s

Treat investigation capacity as a budget and set the bar from cost asymmetry: a missed regression versus a wasted analyst day. Tier the metrics, publish a calendar of known effects, and require a written triage note before escalating.

solid answer

~50 s

I start by naming the tradeoff out loud: a lower bar buys faster detection and costs analyst days, a higher bar buys focus and costs late discovery. Then I make three structural moves. First, tier the metrics: a small set of decision-grade metrics with named owners and pre-agreed thresholds, everything else on dashboards with a review cadence and no alarm. Second, publish a known-effects calendar — releases, holidays, marketing pushes, pipeline changes — so predictable wobbles never trigger a drill in the first place. Third, require a cheap standard triage note before any escalation: the comparison used, how the delta compares with normal variation, and a data-quality check. Most dips die there in ten minutes. I also require persistence for anything ambiguous — two periods or a confirming independent source. Finally I measure the programme: record the outcome of every alarm, and retune the thresholds on the observed hit rate rather than on whoever shouted loudest.

go deeper

for a junior

Understand that an investigation costs real time, and that writing down what you checked and found nothing is a useful result rather than a failure.

for a middle

Be ready to propose a concrete rule — a threshold plus a persistence requirement — and to say what that rule would have flagged over the past year of data.

for a senior

Show you can hold the line on a noisy dip while still giving stakeholders a defensible answer, and that each investigation's outcome feeds back into the threshold.

for a principal

Own the tradeoff explicitly: the false-alarm rate you are buying, the miss rate you accept, which metrics deserve standing alarms, and how you defend that to an executive who wants everything explained.

## The real question 'Which dips do we investigate?' is a resource-allocation decision dressed as a statistical one. Analyst attention is finite, and every investigation has an opportunity cost measured in the work that did not happen. When almost every investigation ends in 'it was within normal variation', the organisation is paying that cost repeatedly for no information — and worse, it is training people to stop taking alarms seriously, which is how a genuine regression gets missed in a stream of noise. ## Frame it as two error costs Every threshold trades two mistakes against each other. Setting the bar low means investigating things that turn out to be nothing: a known, quantifiable cost in analyst days and interrupted roadmap work. Setting it high means real regressions run longer before anyone looks: a cost in lost revenue, damaged trust, or a bug reaching more users. The right bar is not a statistical constant; it depends on which of those hurts more for that particular metric. That immediately implies the bar should differ across metrics. A checkout success rate that can silently lose money justifies a sensitive threshold and a pager. A share-button click rate does not — its regressions are cheap and its variance is high, so a monthly review is the correct response and an alarm is a mistake. ## Move 1 — tier the metrics A short, explicit tiering makes the whole problem tractable: - **Decision-grade metrics** — a handful, each with a named owner, a pre-agreed threshold, a defined action when it fires, and a documented normal variation. If nobody can say what action a breach triggers, it does not belong in this tier. - **Health metrics** — watched on a cadence, reviewed in a weekly meeting, no alarms. Movement is discussed, not escalated. - **Exploratory metrics** — available for analysis, never a trigger for anything. Most organisations get into trouble by treating everything on the main dashboard as tier one by default, so any wobble anywhere can start a fire drill. ## Move 2 — remove the predictable causes before they alarm A large share of investigations resolve to something someone already knew: a holiday, a marketing burst that ended, a release, a pricing change, a pipeline migration, a big client's invoicing cycle. Maintaining a shared, annotated calendar of these events — visible on the dashboards themselves — converts that whole class from investigations into captions. This is often the single highest-return change available, because it costs a habit rather than a headcount. ## Move 3 — make triage cheap and standard Between 'ignore it' and 'open an investigation' there should be a fixed, ten-minute step with a written output: which comparison was used, how the delta compares with the metric's normal variation, whether the data is complete and the instrumentation intact, and whether any known calendar event applies. Most dips die at this stage, and the note is the artefact that stops the same dip from being re-litigated next week. Recording 'checked, inside normal variation' is a result worth keeping, not an admission of failure. For ambiguous cases, add a persistence requirement: two consecutive periods, or confirmation from an independent source, before an investigation is opened. Persistence costs a day of patience and removes a large share of chance findings. ## Move 4 — measure the programme itself Thresholds should be tuned on evidence like anything else. Log every alarm and its eventual resolution — real cause found, known calendar effect, data-quality issue, or nothing — and separately log regressions that were discovered by other means, such as customer complaints or a partner escalation. Those two logs tell you which way to move the bar. If almost everything resolves to nothing, the bar is too low. If problems keep arriving through support tickets before the alarms fire, it is too high. Review this quarterly, not in the middle of an incident. ## Defending the bar upward A senior stakeholder will eventually ask for every 2% dip to be explained. The answer is not a refusal but an accounting: show how often a 2% move happens under normal variation, multiply by the hours each investigation costs, name what would go undone, and offer the tiered alternative — a fast triage note for everything, a full investigation for moves outside the agreed band or sustained across periods. Framed as a budget with a stated miss rate the organisation is choosing to accept, this is a conversation about priorities rather than about analytical rigour, and it is one a lead is expected to run.

  • How do you decide which metrics get a standing alarm at all?
    The ones a real decision hangs on, with a named owner who would take a specific action when it fires. If nobody can say what happens on a breach, the metric needs a review cadence, not an alarm. That test alone removes most of the dashboard from the alerting tier and makes the remaining alarms credible.
  • How do you know whether the bar is set at the right level?
    Log the resolution of every alarm — real cause, known calendar effect, data issue, or nothing — and separately log regressions discovered by other routes such as customer complaints. If nearly all alarms resolve to nothing, the bar is too low; if problems keep arriving through support before the alarms fire, it is too high. Retune quarterly on those two logs.
  • A senior stakeholder wants every 2% dip explained. How do you respond?
    With arithmetic, not resistance. Show how often 2% occurs under normal variation, multiply by hours per investigation, and name the work that would be displaced. Then offer the tiered alternative: a ten-minute triage note for every dip, full investigation reserved for moves outside the agreed band or sustained over two periods.

saying these in an interview costs you the question

  • Investigates every dip because a stakeholder asked
  • Sets thresholds with no estimate of the false-alarm rate
  • Never records the outcome of past investigations
  • Treats alert fatigue as an individual discipline problem
  • Fires alarms with no owner and no agreed action

context