One worker identity in a fleet that reads two values at start-up logged 40,000 reads yesterday — what does that volume establish?
answer
- divide by the reads per start
- three orders of magnitude, not a percentage
- a crash loop reads exactly the same way
- volume says look, not guilty
- check whether a second dimension moved
basics
~20 sOnly that the reads are far beyond any need this job has, and must be explained. A restart loop, a retry storm or a cache that stopped caching produces the same count. Volume says look, not guilty.
solid answer
~40 sStart with the arithmetic the shape makes available. Two reads per start means 40,000 reads implies about 20,000 starts in a day — roughly fourteen a minute, for a fleet that normally deploys a few times a day. That is a departure by three orders of magnitude, so the finding is solid: nothing about this job needs that many reads. What it does not establish is who caused it. The most common cause is a worker crash-looping and re-reading on every start, and a close second is a caller that used to cache the value and stopped. Both are indistinguishable from bulk harvesting on the count alone. The dimensions that separate them are the ones alongside it — whether the source, the hour and the set of names also moved.
go deeper
Know that a big read count is a reason to ask what happened, not proof of an attack. The same number comes out of a service that keeps restarting.
Do the arithmetic. Divide the total by the reads a single start costs, state the implied start count, and compare it with what the fleet actually deploys.
Show how you disambiguate: check whether source, hour or the set of names moved alongside the volume, and say which combination points at a fault versus a credential in the wrong hands.
Decide what a volume departure is worth across the estate — where you want a rate limit, where you want an alert, and who owes an explanation when a count has no owner.
## Do the arithmetic before anything else The fleet's shape gives you a divisor, so use it. Two reads per start and 40,000 reads implies roughly **20,000 starts** in a day. Against a fleet that deploys a handful of times a day and otherwise runs steadily, that is not a percentage departure — it is three orders of magnitude, about fourteen starts per minute sustained. Saying the number out loud matters in an interview, because it changes what you are claiming. "Reads are up" is a hunch. "This identity performed 20,000 start-ups in a day for a job that starts a few times a day" is a statement someone can confirm or refute in a minute, and it is what makes the rest of the conversation concrete. ## What the count establishes, and what it does not **Establishes:** a departure worth explaining. The volume is beyond any need the job has, so either the job is not behaving as designed or the credential is not being used by the job. **Does not establish:** that this is abuse. At least four things produce the same count: - a worker in a **restart loop** — it crashes, is replaced, reads its two values, crashes again; the count is a symptom of a crash, not of theft - a caller that **stopped caching** — a change removed an in-process cache, so a value that was read once per start is now read per unit of work - a **retry storm** against a downstream system, where each retry path re-fetches the credential instead of reusing it - **bulk harvesting** — somebody reading everything they can reach, quickly, before anyone notices Every one of those is composed entirely of permitted reads by a valid credential. None of them registers as a failure anywhere. ## The dimensions that separate them | Also observed | Points toward | |---|---| | Same two names, same source, same hours | an operational fault in the fleet | | Same names, but a source never seen before | the credential in use somewhere the fleet does not run | | Many distinct names, or a first-ever listing | reaching beyond need, not a fault | | Concentrated in a window the service does not run | a caller unconstrained by the service's schedule | | A burst of issuance requests alongside the reads | one identity asking the store to mint far more than it has ever needed | This is the practical reason a baseline is a set of fields rather than a single counter. Volume on its own is ambiguous by construction; volume plus a second moved dimension is a much narrower story. ## The counter you actually want A raw daily total is the wrong shape for this job. Two better ones: 1. **Reads per start, per identity.** For this fleet the expected value is exactly two. A worker reading two values thirty times per start is a different and more interesting finding than a worker starting thirty times. 2. **Reads against the job's stated need.** Recording that the job needs two names lets you express the departure as a ratio rather than a number, which travels across identities with different shapes. Both of those are per identity. An estate-wide reads-per-hour graph would absorb 40,000 extra reads without a visible bump in most estates, which is precisely how this kind of activity stays quiet. ## Why the operational causes are not a reason to ignore it It would be easy to conclude that because restart loops are common and harvesting is rare, a volume signal is mostly noise. Two reasons that conclusion is wrong. First, the common causes are worth knowing about anyway: a crash-looping fleet and a cache that silently stopped are both real production defects, and this signal found them. Second, and more importantly, **the cost of the check is an explanation, and the cost of skipping it is unbounded.** The signal's job is to make somebody say out loud why 40,000 reads happened. If the answer is "the workers were crash-looping between 02:00 and 04:00", that is a good outcome and takes minutes. If nobody can produce an answer, that absence is the finding, and it arrives with the identity, the time window and the names already attached — which is a far better starting position than learning about it from somewhere else weeks later.
- The source, hour and names all match the baseline; only the count moved. What is your leading hypothesis?An operational fault in the fleet — a restart loop or a lost cache. When the only dimension that moved is volume and everything about where, when and what stayed put, the simplest explanation is that the job is doing its normal thing far too often.
- Why is 'reads per start' a better counter than a daily total here?Because it separates two different findings that the total conflates: a worker starting far too often, and a worker re-reading far too much within one start. The first is usually a crash loop; the second means something changed about how the value is used.
- Would a per-identity read rate limit have been better than an alert?It bounds the damage but also masks the signal, and it will eventually refuse a legitimate surge. If you set one, treat hitting it as an event to investigate rather than a problem solved, or you have converted a visible departure into a quiet refusal nobody reads.
saying these in an interview costs you the question
- Declares theft from the read count alone
- Ignores a restart loop as an explanation for high volume
- Compares the count to an estate-wide total instead of the identity's own
- Reports 'reads are up' without dividing by reads per start
- Dismisses the signal as noise because operational causes are common