Field usage analytics reports zero resolutions of Listing.priceHistory for 87 days - what could make that reading wrong?
answer
- A zero is a measurement, interrogate it
- Sampling misses rare, never misses hot
- Cached reads resolve nothing at all
- How big is the unknown-client bucket
- Window versus the slowest client's release cycle
basics
~20 sZero recorded usage proves the collector saw nothing, not that nobody used the field. Sampling, uninstrumented instances, cached responses, unattributed traffic, rare or seasonal callers and a retention window shorter than the claim can all produce a zero for a live field.
solid answer
~50 sTreat a zero as a measurement, and interrogate the measurement first. **Collection gaps**: head sampling at one in 128 makes a field used a few dozen times a quarter almost certainly invisible; one region, replica or older build may not be instrumented; the reporting pipeline may have dropped a window; retention shorter than 87 days silently truncates the claim. **Traffic that never executes**: a response served from an edge or server-side cache resolves nothing. **Attribution gaps**: executions arriving without client-identification headers land in `unknown`, so "no client uses it" is unprovable. **Real but rare usage**: a quarterly reconciliation job, a shipped mobile build that only calls it on a rarely-taken screen, an integration partner who queries at year end. And **structurally hidden usage**: a coordinate whose parent is usually null never resolves, yet clients still select it and their documents would fail validation without it.
code
pseudocode · 12 lines# updated, never appended - bounded by schema size x client builds
upsert last_seen(
coordinate = "Listing.priceHistory",
client_name = "unknown",
client_version= "unknown",
last_at = max(existing.last_at, now())
)
# the deletion question becomes a lookup, not a retention gamble
rows = last_seen.where(coordinate == "Listing.priceHistory")
quiet_for = now() - max(rows.last_at, default = FOREVER)
coverage = attributed_executions / total_executions # publish this toogo deeper
Take away the headline: a zero in a usage report means the collector recorded nothing, which is not the same as nobody using the field. Be able to name one gap, such as sampling or a cached response.
Explain the mechanisms behind the gaps - how head sampling erases rare events, why a cache hit resolves no fields, and why unattributed requests make a per-client claim unprovable.
Demonstrate the operating judgement: check sampling rate, instrumentation coverage, retention, attribution coverage and cache hit rate before quoting the window, and keep an alert on the coordinate so a resurrection is caught while it is still reversible.
Own the standard of evidence. Decide what the organisation may conclude from a quiet coordinate, what it must capture unsampled to be allowed to conclude it, and how that standard changes for a graph whose callers you cannot see.
## Zero is a measurement, not a fact The sentence "this field has had no usage for 87 days" contains a hidden claim: that the collector would have seen the usage if it had happened. Interrogating that claim is the whole senior skill here. There are five families of reason a live field reads as zero. ## 1. The collector did not see it **Sampling** is the biggest one, and the one people turn on for cost reasons without thinking about this consequence. Sampling is safe for "how hot is this coordinate" and actively harmful for "is this coordinate dead", because the error is asymmetric: a hot field survives any sampling rate, while a rare field disappears. At one-in-128 head sampling, a field genuinely resolved 40 times in the window has roughly a 73% chance of never once being sampled. If you must sample, sample the volume and timing stream and never the coordinate-set stream. **Partial instrumentation** is the quiet one. One region, one canary deployment, one old build still running, or one internal replica that predates the hook. Also: an ingestion pipeline that dropped a day, or a retention policy shorter than the window you are quoting, which turns "no usage in 87 days" into "no usage in the 30 days we still store". ## 2. The request never executed the field A response served from a server-side response cache or an edge cache resolves nothing at all: the client got the data, the coordinate recorded no usage. A client-side store answering a screen from memory produces no request whatsoever. In both cases the field is genuinely load-bearing and genuinely invisible. Any coordinate that sits under a heavily cached read path needs the cache's own hit metrics read alongside the usage data. ## 3. The execution happened but could not be attributed Client name and version arrive by header convention, and both are self-reported. Requests that carry neither - a script, a partner, a service-to-service caller, an internal tool - collapse into `unknown`. The useful discipline is to publish the **attribution coverage** number alongside every usage report: if 18% of executions are unattributed, then "no client uses this" is a statement about 82% of your traffic, and the honest phrasing says so. ## 4. The usage is real but rare This is where domain knowledge beats data. In a real-estate listings graph, `Listing.priceHistory` might be read only by a price-correction screen an agent opens when a listing is disputed, or by a quarterly market-report job, or by a partner feed pulled at the end of a financial year. 87 days is a long window for a web client that ships weekly and a very short one for a mobile binary at version 6.14.2 that a slice of users will not update for a year, or for a business process with an annual cycle. The practical rule: the window has to exceed both the slowest client's release-plus-adoption cycle **and** the longest business cycle that touches the data. Neither number comes from the usage system. ## 5. The usage is structurally hidden A coordinate under a parent that resolves to null for almost all data never resolves. Execution-derived usage reads zero, while documents in the field still select it - and those documents fail validation the moment the field leaves the schema. This is the case where an otherwise perfect zero is most likely to break someone, and the mitigation is to keep the coordinates that *validated documents selected* as a second signal beside the coordinates that *executed*. ## Making a zero worth trusting The things that turn a weak zero into a strong one are all decided before the question is asked: - **Never sample the coordinate-set stream.** Full capture of the distinct coordinates per execution, sampling only for volume and latency. - **Store a last-seen timestamp per coordinate, per client, per version**, and keep it indefinitely. It is tiny - bounded by schema size times client builds - and it makes the window a query parameter instead of a retention limit. - **Measure and publish attribution coverage**, so the unknown bucket is visible rather than assumed empty. - **Read cache hit rates for the same path** before concluding a cached field is unused. - **Alarm on resurrection**: keep the instrumentation reporting on a coordinate you believe is dead, so that if it resolves once after the decision, someone hears about it while the change is still reversible. What you can honestly conclude is bounded: "no execution we recorded, from any client we could identify, resolved this coordinate in 87 days, and we capture the coordinate set on 100% of executions with 94% attribution coverage." That is a strong statement. "Nobody uses it" is not the same statement, and the difference is where the outage comes from.
- Why is sampling acceptable for latency data but not for this?Because the error is asymmetric. A sample estimates a distribution well when the population is large, which is exactly the case for a hot field's latency. The deletion question is about the rare tail: a coordinate resolved a few dozen times in a quarter is precisely the one sampling erases. Sample the volume and timing stream freely; capture the distinct-coordinate set on every execution.
- How long should the observation window be?Longer than both the slowest client's release-plus-adoption cycle and the longest business cycle that touches the data. A web client shipping weekly is satisfied by weeks; a shipped mobile build with a long tail of un-upgraded installs and a partner integration that runs at financial year end are not. Neither number comes out of the usage system, which is why the window is a judgement call informed by client telemetry.
- How would you notice that a coordinate you believed was dead came back?Keep collecting on it and alert on the first resolution. A last-seen row that updates after the coordinate was declared quiet is the cheapest possible signal, and it arrives while the change is still reversible. The alert should carry the client name and version so the follow-up is a conversation with an owner rather than an investigation.
- A field is heavily cached at the edge. How do you reason about its usage data?Read the cache's hit rate for that path alongside the usage record. A high hit rate means most reads never reached execution, so the recorded resolutions are a lower bound scaled by the miss rate rather than a measure of demand. If the read path is cached aggressively enough, usage analytics alone cannot support a deletion argument for anything underneath it.
An empty footfall count for a gallery is only evidence the room is unused if the sensor was switched on, in every entrance, for the whole period - and nobody was watching the exhibit through the window.
saying these in an interview costs you the question
- Treats zero recorded usage as proof of no usage
- Leaves sampling on for the coordinate-set stream
- Quotes a window longer than the data's retention
- Ignores the unattributed-client bucket entirely
- Forgets that cached responses resolve no fields
- Assumes all client builds update within weeks