Each of 4,000 issuance calls from one identity was authenticated, authorized and logged, yet nothing alerted — what must be measured?
answer
- the stolen credential succeeds
- failure alerting sees nothing
- count, do not inspect
- per identity, not per fleet
- 300x on one, 37% on the total
basics
~20 sIssuance volume per requesting identity, against that identity's own baseline. Failure-shaped alerting stays silent because every call succeeded, and no single request in the loop is anomalous — the count in a window is the only thing that moved.
solid answer
~40 sThe default alerting in most estates is failure-shaped: authentication failures, denied requests, error rates. A borrowed identity generates none of those, because it is an allowed caller doing an allowed thing with a valid credential. What moved is the **count**. Three counters catch it: mints per identity per interval compared with that identity's own baseline; live credentials per identity compared with what that identity legitimately needs; and a fleet-wide issuance rate, which catches many identities each staying under their own threshold. The per-identity baseline is the load-bearing one — a collector whose normal is 2 mints an hour going to 600 is a 300-fold change, while the same 600 added to a fleet baseline of 1,600 an hour is a 37% bump that a global threshold will not see.
go deeper
Recall the core asymmetry: a stolen credential succeeds, so alerting built around failures and denials cannot see it.
Explain which counters move — mints per identity per interval, live credentials per identity, and the fleet total — and what each one catches that the others do not.
Show that you would disaggregate by requesting identity rather than alert on a fleet total, put the responder's first questions inside the alert payload, and handle deploy windows before the rule is muted.
Own the detection story's boundaries: state plainly that issuance-volume alerting covers minting and not misuse of already-issued material, and decide who owns the baselines as the fleet changes shape.
## Why the alerts that exist stayed quiet Walk through what the loop actually produced. Every call arrived with a valid workload identity. Every call was permitted by policy — this identity is supposed to request credentials for this target. Every call succeeded and was written to the request log. So the estate's standard security signals all read normal: - **Authentication failures**: zero. The attacker had a working identity. - **Denied requests**: zero. Nothing in the policy said no. - **Error rate**: unchanged. The store answered every call. - **Latency and capacity**: unchanged. The volume was a small fraction of fleet throughput. This is the general shape of credential abuse: the stolen thing **succeeds**. Alerting built around failure is blind to it, and saying "our failed-authentication alert would have caught this" is the wrong answer to this question. ## The measurement that does move The only thing that changed is how many credentials came out, so that is what has to be counted. Three counters, each catching something the others miss: 1. **Mints per identity per interval, against that identity's own baseline.** The primary signal. Workload identities are extraordinarily regular: a collector that mints twice an hour does so every hour, forever. 2. **Live credentials per identity, against what that identity legitimately needs.** This needs no history at all. A collector that should hold one credential holding four hundred is an alert on the ratio alone, and it also catches a slow accumulation that never trips a rate threshold. 3. **Fleet-wide issuance rate.** The backstop for the spread case — many identities each used gently, every one of them under its own ceiling, with only the total showing the campaign. ## Why the per-identity baseline is the load-bearing one Numbers make this concrete. Assume 800 collectors, each legitimately minting 2 credentials an hour: a fleet baseline of **1,600 an hour**. Now a single compromised identity mints at 10 a minute — **600 an hour**: | Viewed as | Baseline | With the attacker | Change | |---|---|---|---| | One identity's rate | 2 / hour | 602 / hour | about **300x** | | Fleet-wide rate | 1,600 / hour | 2,200 / hour | about **+37%** | A fleet-wide threshold set at, say, double the baseline never fires. The same traffic seen against the one identity that produced it is unmissable. Aggregation is what hid the attack, and disaggregating by the requesting identity is what reveals it. ## What the alert has to carry An alert that says "unusual issuance volume" and nothing else costs the responder the first ten minutes of the incident, and in this attack ten minutes is 2,000 more credentials. Put the answers in the alert: - the **requesting identity** and the session it was minting under; - the **count in the window** and the baseline it is being compared against; - the **targets** the credentials were minted against, so the reach is visible immediately; - the **earliest and latest expiry** among them, which is how long the tail runs if nothing is withdrawn; - whether a cap fired, and which one. ## Thresholds that survive contact with a real fleet - **Deploys and scale-outs look like attacks.** A rollout restarts hundreds of consumers, each of which mints on start-up. Either the alert knows about change windows, or its first month is all false positives and it gets muted. - **Absolute floors beat pure ratios for quiet identities.** An identity whose baseline is 0.1 mints an hour trips a multiplier on any activity at all. Require both a ratio and a minimum absolute count. - **Alert on the sustained window, page on the burst.** A count over five minutes catches the loop while it is still running; a count over a day catches the patient version. Both are worth having and they are not the same rule. ## The boundary of this signal This measures **minting**. It is silent about a credential that was minted legitimately and is now being used by someone other than the consumer it was issued to — that is a signal about how issued material is being used, and it is read on a different surface with different evidence. Keep the two separate when you describe your detection story, because an interviewer will notice if you claim one covers the other.
- Why is a fleet-wide issuance threshold not enough on its own?Because aggregation hides a single identity inside normal fleet noise. With 800 identities minting 2 an hour, one compromised identity adding 600 an hour raises the total by about 37% — inside ordinary variation, and far below any threshold you could set without constant false alarms. The same traffic is a 300-fold change against that one identity's own baseline. Keep the fleet rate as a backstop for spread campaigns, not as the primary signal.
- Which signal needs no historical baseline at all?The ratio of live credentials to legitimate need. If an identity class is declared to hold at most a handful of credentials at once, then any member holding hundreds is anomalous on its own terms, with no history required. It is also the signal that survives a slow campaign, because an attacker minting under the rate threshold still accumulates a live set that the declared need cannot explain.
- How do you keep a fleet rollout from drowning this alert in false positives?Give the rule a change-window input, since a rollout legitimately restarts hundreds of consumers that each mint on start-up. Pair the ratio with an absolute floor so quiet identities do not trip on any activity, and split the rule into a short window that pages on a burst and a long window that opens a ticket on slow accumulation. An alert that is muted in its first month protects nothing.
saying these in an interview costs you the question
- Failed-authentication alerting would have caught this
- A fleet-wide issuance threshold is sufficient
- Inspect individual requests for something anomalous
- Access logs alone constitute detection
- Alerting on issuance duplicates use-side anomaly signals