skip to content

Diagnosing a Metric Change

Telling a real metric movement from noise: control-chart bands around a week-over-week delta, day-of-week effects, and logging breaks that fake a drop. Nearly every analyst screen opens here.

on this pageshow

questions

6

Daily active users are down 8% versus last Tuesday — what do you check before calling it a real drop?

level: juniorimportance: must knowfreq 82%

answer

  1. How big is a normal week?
  2. Compare like with like
  3. Same weekday, check the calendar
  4. Spread of past week-over-week changes
  5. Broken data before broken product

basics

~20 s

Check whether 8% is inside the normal spread of Tuesday-over-Tuesday changes for this metric, then rule out calendar effects and a tracking break. Only a delta outside normal variation, measured on clean data, is worth investigating as a product problem.

solid answer

~50 s

I start by asking how big a normal week is. I pull the last several same-weekday comparisons — each Tuesday against the prior Tuesday over recent months — and see where -8% falls among those historical changes. If week-over-week swings routinely run plus or minus 6-10%, an 8% move is unremarkable and I say so. Second, I check the calendar: was either week a public holiday, a promotion, or an outage week that makes the comparison invalid? Comparing like with like matters because most product metrics carry a strong day-of-week pattern. Third, I question the data before the product: did a release, a client change, or a pipeline delay change how the metric is collected? Only when the delta survives all three — larger than normal variation, a clean calendar comparison, trustworthy instrumentation — do I start hunting for a product or market cause.

go deeper

for a junior

Be ready to name the first three checks out loud: is the comparison like-for-like, is the delta bigger than normal week-to-week swings, and is the data itself intact.

for a middle

Explain how you would actually measure normal variation — the distribution of past same-weekday changes — and why holiday or promotion weeks must be excluded from that baseline.

for a senior

Show triage discipline under pressure: what you check first, what you refuse to escalate, and how you communicate an inconclusive read to a nervous stakeholder.

for a principal

Own what the organisation should do with dips of this size by default — which metrics get a standing rule, and what an 8% wobble costs in analyst time when everyone chases it.

## The question behind the question When a stakeholder says a metric is down 8%, they are asserting three things at once: that the number moved, that the movement is unusual, and that something caused it. Only the first is established by the dashboard. The job of triage is to test the other two cheaply, before anyone spends a day building a narrative. ## Step 1 — how big is a normal week? A percentage change means nothing without the distribution it came from. The cheapest baseline is the metric's own history of comparable changes: for each of the last 12-20 Tuesdays, compute the percent change against the Tuesday before it. That gives an empirical picture of ordinary movement. If those historical changes commonly land between -9% and +11%, then -8% is an ordinary Tuesday and there is nothing to explain. If they almost never leave plus or minus 2%, then -8% is genuinely striking and deserves attention. Two practical cautions. First, exclude periods you already know were abnormal — an incident week, a launch week, a holiday week — from the baseline, because they inflate the apparent spread and make the band so wide that real drops hide inside it. Second, note the direction of the metric's growth: for a series that grows a few percent a week, a flat week is already a small negative surprise, and the baseline of changes should be centred on that growth rate, not on zero. ## Step 2 — is the comparison like-for-like? Most consumer and B2B product metrics have a pronounced day-of-week shape: weekday and weekend behaviour differ enough that comparing Tuesday with Sunday produces a double-digit 'change' with no cause at all. That is why the default comparison is same weekday against same weekday, or a seven-day total against the previous seven-day total, which averages the weekly pattern away. The calendar breaks even that discipline. A public holiday, a shopping event, a school term boundary, a major sports final, or a marketing push in one of the two weeks makes the pair incomparable. Two honest ways out: compare the affected week with the same calendar event a year earlier rather than with the adjacent week, or annotate and skip it and read the trend around it. What you should never do is quietly compare a holiday week with a normal week and report the gap as a regression. Deeper machinery for modelling repeating seasonal structure belongs to time-series work; at triage level, the requirement is only that the two things you subtract are comparable. ## Step 3 — did the product move, or the measurement? Before accepting that user behaviour changed, ask whether the way the metric is produced changed. A client release, a tagging change, a definition edit, a warehouse job that ran late, a filter that started dropping traffic — any of these can create a clean, convincing drop in a metric while the product is untouched. The fastest test is reconciliation against a second source that the product itself writes. Instrumentation failure is common enough that it belongs in the first three checks, not in the last three. ## Step 4 — only now, look for a cause If the delta is larger than the metric's normal variation, the comparison is like-for-like, and the data pipeline is healthy, then you have a real change and can start looking for a mechanism: a release, a pricing or ranking change, an upstream partner, a competitor, a seasonal demand shift, or a specific segment. That is also the point at which breaking the delta down becomes worthwhile — and where discipline is needed, because a wide-open search across many slices will always surface something alarming. ## How to report an inconclusive triage Saying 'this is inside normal variation' is a result, not a failure, and it should be reported with the evidence attached: the range of historical same-weekday changes, the calendar note, and the data-quality check. Written down, it stops the same dip from being re-litigated next week and builds the shared intuition of how noisy each metric actually is. The strongest analysts are recognised less by the causes they find than by the false alarms they defuse in ten minutes with a baseline.

  • How would you quantify normal week-to-week variation from a year of daily data?
    Build the series of same-weekday percent changes — each Tuesday against the prior Tuesday — for the last 12 to 20 comparisons, and look at their spread and extremes. If -8% sits comfortably inside what those changes normally do, it is noise. Drop weeks you already know were abnormal, such as holidays or incident weeks, or they inflate the spread and hide genuine drops.
  • Why compare a Tuesday with the previous Tuesday rather than with Monday?
    Because most product metrics have a strong day-of-week pattern: weekday and weekend levels can differ by tens of percent for reasons that have nothing to do with product health. Comparing the same weekday holds that pattern constant, so any remaining difference is a candidate for a real explanation.
  • The drop is real and started sharply on a single day. What does that timing tell you?
    A sharp onset points at a discrete event rather than a gradual behavioural shift: a release, a config or pricing change, a partner or channel outage, a policy change. Line the exact hour of the break up against the release and incident timeline, and check whether the drop is confined to whatever that change touched.

An 8% drop in a car's fuel range is only alarming once you know the range normally varies by 10% with weather, traffic and load.

saying these in an interview costs you the question

  • Treats any week-over-week drop as a regression without a baseline
  • Compares against yesterday and ignores the day-of-week pattern
  • Never asks whether the tracking itself changed
  • Calls 8% significant with no idea of normal variation
  • Starts slicing segments before validating that the number is real

context

open as a page

How do you tell a broken analytics pipeline from a genuine drop in signups?

level: middleimportance: must knowfreq 70%

basics

~20 s

Reconcile the metric against an independent source of truth, such as the records the product itself writes. A tracking break usually starts abruptly at a release, hits only one platform or app version, and leaves the downstream business records unchanged.

open as a page

How would you set a control band on weekly revenue to decide which weeks to investigate?

level: middleimportance: should knowfreq 45%

basics

~20 s

Pick a baseline of stable weeks, take a centre line from them, and estimate the spread from consecutive-week differences. Set limits at the centre plus and minus about three spreads; weeks outside get investigated. Exclude known-abnormal weeks from the baseline.

open as a page

You slice a 3% drop across 40 country-by-platform cells and one is down 30% — what now?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Ask how much of the aggregate drop that cell can explain, and whether an extreme cell was expected anyway. Scanning 40 slices guarantees a few look alarming, and small cells swing hardest, so size the contribution first.

open as a page

Total revenue fell 6% while paying-customer count is flat — how does account concentration change your read?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Revenue is usually concentrated in a few large accounts, so check whether one or two explain the whole 6%. Recompute the delta excluding the top accounts: if it vanishes, this is an account event, not a broad product change.

open as a page

Most metric-dip investigations on your team find nothing — how do you set the bar for investigating?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Treat investigation capacity as a budget and set the bar from cost asymmetry: a missed regression versus a wasted analyst day. Tier the metrics, publish a calendar of known effects, and require a written triage note before escalating.

open as a page