Your dashboards are flat and green during a provider incident, yet customers report errors and the provider's status page says all clear — which signal do you trust?
answer
- flat is not the same as zero
- the instrument shares the bloodstream
- absence of data is itself a signal
- probe from outside the platform
- status pages are human-gated and lag
basics
~20 sTrust the customers. Telemetry ingestion is itself a shared platform service, so a flat chart is missing data rather than healthy traffic, and a status page is gated on human confirmation. The signal to believe is the one sharing fewest dependencies with the thing it measures.
solid answer
~50 sA flat line has two readings — *the value is zero* and *no value arrived* — and they render almost identically. During a platform-wide incident the second is far more likely, because the ingestion path that carries your measurements is itself a service the provider operates and every tenant shares. The alerting built on top inherits that: alerts that fire on a threshold cannot fire on data that never arrives. The status page is not a substitute, because publishing it is gated on the provider detecting, confirming and scoping the impact first. So rank your signals by **independence from the failure**: customer reports and a probe running outside the platform are worth more right now than any chart inside it. The practical rule is that silence from a monitoring system is not evidence of health.
go deeper
Learn to ask what a flat chart means before reacting to it: no data at all looks almost the same as a value of zero. When a customer says something is broken and the chart says nothing is, the customer is the stronger evidence.
Explain the path a measurement travels — collect, authenticate, resolve, send, store, query — and name which steps a platform incident can break. That chain is why alerting inherits the outage and why silence is not health.
Show the practical kit: a heartbeat series, a staleness alarm, an out-of-band probe sharing no platform dependency, and alert delivery that does not run on the affected platform. Rank signals by independence, not by richness.
Decide how much the organisation pays for independent observation, and set the standard that at least one detection path must not share dependencies with what it watches. Say honestly which outages you have chosen to be blind to.
## Two readings of a flat line Every time-series chart draws a gap and a zero almost identically, and the difference between them is the whole question. *The value is zero* means the measurement arrived and said nothing happened. *No value arrived* means the pipeline that carries measurements is broken and the chart is showing you its own silence. In normal operation the first reading is usually right, which is exactly why teams reach for it reflexively at the worst possible moment. During a platform-wide provider incident the second reading is far more likely, for a structural reason: **the telemetry ingestion path is itself a shared platform service**. It is operated by the provider, used by every tenant, and it shares the same underlying identity, name-resolution and network dependencies as the workloads it is measuring. It is not an independent observer. It is a participant in the same failure. ## Why the monitoring shares the failure Work through what a measurement has to do to reach a chart. An agent in the workload collects it, authenticates to the telemetry service with a platform credential, resolves that service's name, sends it over the platform network, and the service stores it for a query to find later. Every one of those steps is a dependency the incident may already have taken out. - If the platform's identity service is unavailable, the agent cannot authenticate and the measurement is dropped at the source. - If name resolution is degraded, the agent cannot find the ingestion endpoint. - If ingestion is degraded, the data is accepted and lost, or rejected outright. - If the query side is degraded, the data is safely stored and you cannot read it. The alerting built on those series inherits the whole chain. A rule that fires when errors exceed a threshold cannot fire on data that never arrived — its silence looks exactly like health. This is the single most dangerous property of the failure class: the instrument and the patient share a bloodstream. ## Ranking signals by independence The useful question in the first ten minutes is not *what does each signal say* but *what does each signal depend on*. | Signal | Depends on | What it can still tell you | |---|---|---| | Platform dashboards and alerts | The provider's telemetry ingestion, identity and network | Very little; its silence is uninformative | | Provider status page | The provider detecting, confirming and scoping first | Scope and duration, late | | Customer reports and support volume | Nothing you operate | That impact is real, with no detail | | A probe running outside the platform | Its own separate host and network | Whether your endpoint answers, right now | That ordering inverts the everyday one, and that inversion is the answer to the question. An engineer who has lived through this reaches for the least sophisticated signal first, because sophistication here is bought with dependencies. ## Telling blindness from health before it matters The distinguishing work is cheap and has to be done in advance: 1. **Emit a heartbeat series that is never zero.** A constant, always-present counter turns 'no data' into a detectable condition, because a gap in it cannot be explained as quiet traffic. 2. **Alert on staleness, not only on thresholds.** An alarm on *time since the last measurement arrived* fires precisely in the case a threshold alarm cannot cover. 3. **Run at least one probe outside the platform.** It must share no identity, no name resolution and no telemetry path with the target — otherwise it goes quiet in the same event and tells you nothing. 4. **Route the alert out of band.** An alerting channel that itself runs on the affected platform will deliver the page after the incident is over. A probe that only checks whether the endpoint answers, from somewhere entirely separate, is worth more during this incident than a wall of richly dimensioned charts, because it is the only thing in the room that is not also broken. ## What the green status page is worth Not much, in this window. Publishing is gated on the provider detecting the impact, confirming it, scoping which services and which regions are affected, and getting the post approved. The lag is structural rather than negligence, and pages are published per service and per region rather than per tenant, so a degradation touching a subset of tenants may never appear at all. The correct conclusion from a green page thirty minutes in is *almost nothing yet* — never *therefore we are fine*. Treat your own measured customer impact as the trigger, and use the page later for what it is genuinely good at: telling you the scope and the all-clear.
- What makes a flat series distinguishable from genuinely zero traffic?A heartbeat: a series the workload emits unconditionally, at a constant rate, whatever the traffic. Real quiet still produces heartbeats, so a gap in them can only mean the pipeline stopped. Pair it with an alarm on the age of the most recent measurement rather than on its value, and you have an alert that fires on absence — which no threshold rule can do.
- Where should the probe that checks your service run, and what must it not share?Outside the platform entirely, and ideally outside the region under suspicion. It must not use platform credentials, platform name resolution, platform telemetry or the platform's alert delivery, because each shared dependency is a way for it to fall silent in the same event. Keeping it deliberately simple — does the endpoint answer — is a feature, since every capability you add to it is another dependency.
- If dashboards are blind, how do you get any detail at all about what is failing?From the workloads themselves rather than from the aggregation layer: local log files still being written on the instance, an in-process diagnostics endpoint the probe can read, and the error shapes your clients are seeing. It is slower and unsampled, but it is available. The lesson worth carrying out of the incident is that at least one detail path should not route through the platform's telemetry service.
A smoke detector wired into the same circuit as the appliance it watches. When the circuit fails the detector goes quiet, and its silence is not evidence that there is no fire.
saying these in an interview costs you the question
- Reads a flat line as a healthy line rather than as missing data
- Treats a green provider status page as authoritative about your own users
- Assumes alerts would have fired if anything were genuinely broken
- Runs the out-of-band probe on the very platform it is watching
- Cannot distinguish a collector outage from a real drop in traffic
- Takes agreement between two charts fed by one pipeline as corroboration