skip to content

How do you tell a broken analytics pipeline from a genuine drop in signups?

level: middleimportance: must knowfreq 70%

answer

  1. Trust, then verify the number
  2. Is there a second source of truth?
  3. Where exactly does the drop start?
  4. One platform, one app version?
  5. Late-arriving data undercounts today

basics

~20 s

Reconcile the metric against an independent source of truth, such as the records the product itself writes. A tracking break usually starts abruptly at a release, hits only one platform or app version, and leaves the downstream business records unchanged.

solid answer

~60 s

The cheapest test is reconciliation. Signups also exist as things the product writes down — account rows, confirmation emails, first sessions — so if the event stream says signups fell 20% and the account records say they did not, the product is fine and the logging is broken. The shape of the change is the second clue. Tracking breaks are usually step changes that land exactly at a release boundary, are confined to the surface that shipped — one platform, one app version, one web bundle — and often knock out several correlated events together, because the same client code emits them all. A genuine behavioural drop tends to be smoother, spread across platforms, and visible in the downstream funnel too. I also check freshness: late-arriving data makes the most recent day look artificially low, which is the single most common false alarm. The canonical case is an Android release that broke the analytics SDK — signups 'fell' only on new Android versions, and only in the event stream.

go deeper

for a junior

Know that a metric can move because the measurement broke rather than the product, and be able to name one independent source you could reconcile the number against.

for a middle

Explain the signatures of a tracking break — a step change at a release, confined to one platform or app version, several correlated events falling together — and why late-arriving data fakes a drop.

for a senior

Show that you validate data health before opening an incident, and that you can quantify how much of the delta the instrumentation gap explains instead of waving it away.

for a principal

Argue for treating instrumentation as a monitored surface with its own ownership, release gating and freshness alerting, and justify that cost against the analyst time and false incidents it prevents.

## Why this check comes early Every metric is the output of a measurement chain: client code fires an event, a collector accepts it, a pipeline lands it, a query defines the metric. A break anywhere in that chain looks exactly like a change in user behaviour on the dashboard. Because measurement breaks are common — every release touches client code, every schema change touches the pipeline — and because they are cheap to rule out, data validation belongs near the front of any metric-change investigation, not at the end after a week of theories. ## Test 1 — reconcile against a source of record Draw a distinction between telemetry and records of record. Telemetry is best-effort: client events sent over unreliable networks, subject to ad blockers, app crashes, consent banners and SDK bugs. A record of record is written by the system that performs the action: the account row created at signup, the payment row, the order, the ticket. These rarely disappear silently because business processes depend on them. So the first question is: does the same event exist in a second system, and did that second system see a drop? If accounts created is flat while 'signup' events fell 20%, the answer is already settled. It also pays to know the normal gap between the two counts — telemetry typically undercounts by some stable percentage — because a moving gap is itself an early warning signal. ## Test 2 — read the shape and the boundary of the change Instrumentation failures have a recognisable fingerprint: - **Step, not slope.** The series goes from normal to broken between two adjacent buckets, rather than sliding over days. Real behaviour rarely changes discontinuously unless a discrete event caused it. - **Aligned with a deploy.** The break lands at the hour a release rolled out, or tracks the release's staged rollout curve as adoption climbs. - **Confined to a surface.** Split by platform, app version, browser, or client bundle. A break usually lives entirely inside the versions that shipped the bad code, while older versions carry on reporting normally. A behavioural change rarely respects app-version boundaries so exactly. - **Correlated events fall together.** If a screen-view event, a button-tap event and a signup event all fall by a similar proportion on the same surface, one piece of client code stopped firing, because real users do not abandon three unrelated actions in identical proportion. - **Impossible values.** Exact zeros, counts that flatline, duplicated events, or a rate that jumps to a suspiciously round number are measurement artefacts, not behaviour. A genuine drop tends to show the opposite signature: it appears on several platforms, it propagates to downstream funnel steps and to the revenue that follows them, and its size varies across segments rather than being a uniform proportional cut on one surface. ## Test 3 — is the data even complete? The most frequent false alarm has no bug behind it at all: the current day, or the most recent hour, is still filling. Batch jobs land late, mobile clients upload events with delay after coming back online, and a backfill may still be running. Any dashboard read before the data is complete shows an artificial drop that heals itself overnight. The rules that prevent this are simple: mark the newest bucket as provisional, compare only complete periods, and check the pipeline's freshness or row-count monitors before reacting to anything. ## Attribution when both are true Sometimes a release both broke tracking and hurt the product. The way through is to quantify: measure the metric only on the surfaces the release did not touch, or only on the source of record, and see how much of the delta survives. If signups on web are flat and only new Android versions dropped in telemetry while Android account rows held steady, the instrumentation explains all of it. If Android account rows fell by half of the telemetry drop, then both are happening and only half the delta is real. ## Preventing the next one Instrumentation deserves monitoring in its own right: event volume by platform and app version, events per session, the standing gap between telemetry counts and source-of-record counts, and schema validation on arrival. Release annotations on dashboards turn a mystery into a five-second read. The organisational version of the same idea is a tracking smoke test in the release checklist, so a broken SDK is caught by the pipeline that ships it rather than by an analyst three days later, halfway through an investigation into a drop that never happened.

  • The event stream and the account records disagree by 20%. Which do you trust?
    The system that performs the action, not the one describing it. Account rows, payments and orders exist because a business process needs them; client telemetry is best-effort and lossy. Use the source of record for counts and telemetry for behaviour, and track the normal gap between them so you notice when the gap itself moves.
  • Signups look down 15% today but normal by tomorrow morning. What happened?
    Almost certainly incomplete data: the current day was still loading when you looked, or mobile clients uploaded their events late. Treat the newest bucket as provisional, compare only complete periods, and check freshness monitors before reacting. This is the most common cause of an investigation into a drop that never existed.
  • How do you stop this class of false alarm from recurring?
    Monitor the instrumentation as a product surface: event volume per platform and app version, events per session, and the standing gap between telemetry and source-of-record counts. Add a tracking smoke test to the release checklist and annotate release times on dashboards, so a broken client is caught by the release process rather than by an analyst days later.

saying these in an interview costs you the question

  • Assumes the warehouse number is always correct
  • Never checks the app-version or platform breakdown
  • Reacts to today's partial, still-loading data
  • Cannot name a second source to reconcile against
  • Blames the product before checking the release timeline

context