skip to content

Your secret store alerts only on failed authentication, yet a stolen worker credential is being used daily — why does nothing fire?

level: middleimportance: must knowfreq 62%

answer

  1. the read was allowed
  2. failure alerts watch the wrong outcome
  3. abuse here is made of successes
  4. compare a read to that identity's shape
  5. caller, source, hour, breadth, volume

basics

~20 s

A stolen credential authenticates successfully, so failure-based alerting never sees it. The evidence lives in reads that were permitted — an unfamiliar caller, an odd source or hour, more names or more volume than the job needs.

solid answer

~40 s

A failure alert tests the outcome field, and this abuse never sets it. The credential is valid, the policy allows the read, the store serves the value, and the event looks exactly like the one the real worker writes at every start. Denial events still matter — they show a caller probing names it was never granted — but they are the wrong place to look for a credential that is working. The signals that remain all sit on the allowed side: a read from an identity or a source you have not seen before, at an hour the service does not run, across more names than the job needs, or at a volume no consumer could use. Each of those needs a recorded shape of normal reads to be a departure from.

code

pseudocode · 23 lines
pseudocode
// rule A: what a failure-based alert evaluates
on readEvent:
    if readEvent.outcome == DENIED
       and countDenied(readEvent.identity, lastHour) > 5:
        raise "possible credential guessing"

// a stolen but valid credential always sets outcome = ALLOWED,
// so rule A's guard is false on every one of its reads and the
// rule never raises.

// rule B: what a departure-based alert evaluates
on readEvent:
    if readEvent.outcome != ALLOWED:
        return                      // rule B ignores refusals
    shape = baselineFor(readEvent.identity)
    if readEvent.source not in shape.sourceRanges:
        raise "read from an unfamiliar source"
    if readEvent.hour not in shape.hoursSeen:
        raise "read at an unusual hour"
    if readEvent.name not in shape.namesRead:
        raise "read of a name outside this identity's need"
    if countToday(readEvent.identity) > shape.readsPerDayCeiling:
        raise "read volume beyond need"

go deeper

for a junior

Remember the one sentence: a stolen credential is a valid credential, so the read succeeds and a failure alert has nothing to fire on.

for a middle

Explain the mechanics. Name the outcome field the failure rule tests, say why abuse never sets it, and list the dimensions on the allowed side that can show a departure instead.

for a senior

Show you have operated this. Say what has to be recorded per identity before an incident, and be honest that each signal has an innocent twin, so a departure demands an explanation rather than declaring a breach.

for a principal

The tradeoff is coverage against noise across an estate. Decide how much per-identity shape you are willing to maintain, and which departures are worth waking someone for versus reviewing in daylight.

## The two outcomes a store records Every call to a secret store ends one of two ways: the store served the value, or it refused. Almost all first-generation alerting is built on the refusal, because refusals are rare, cheap to count and obviously suspicious. A stolen credential produces the other outcome. It is a **valid** credential — the store has no way to know that the hands holding it changed — so the caller authenticates, the policy check passes, the value comes back, and the event written to the record is identical in outcome to the read the real worker makes at every start. That is the whole of the blind spot: **the abuse you care about is made of successes.** A rule whose trigger is `outcome == denied` tests a field this abuse never sets, so it can run for years and stay silent while the value is read every day. ## What denial events are still good for A candidate who concludes that denials are worthless has over-corrected. They are genuinely useful, just for different things: - a caller reaching for names it was never granted — someone exploring the edges of a grant - a consumer still presenting a value that has been withdrawn, which is usually an operational miss rather than an attack - a credential presented against the wrong branch of the name space, often a deployment error - brute-force noise against the human path, which is loud and low-value but cheap to watch None of those is the case in front of you. Every one of them is a refusal, and your attacker never triggers a refusal because they hold something the store honours. ## The signals that sit on the allowed side Take a concrete shape to depart from: a fleet of message consumers where every worker reads exactly two values at start-up — a queue password and a downstream token — and nothing else for the rest of its life. Against that, the observable departures are: - **an unfamiliar caller** — an identity that has never read this name reads it - **an unfamiliar source** — the same identity, reading from somewhere the fleet does not run - **an unusual hour** — a 3am read on a service that only ever deploys at midday - **breadth beyond need** — an identity that has only ever read two names now walks a whole branch, or issues a list call where it has never listed before - **volume beyond need** — a read count that no consumer of this shape could use - **a burst of issuance** — one identity suddenly asking the store to mint far more downstream credentials than it has ever needed | Signal | Departure from | Also produced legitimately by | |---|---|---| | Unfamiliar caller | which identities read this name | a new service being onboarded | | Unfamiliar source | where this identity has read from | a fleet moved to new address space | | Unusual hour | when this identity runs | an out-of-hours incident or backfill | | Breadth | how many names this identity needs | a genuinely widened job | | Volume | how many reads a start costs | a restart loop or a cache that stopped caching | The right-hand column is why none of these is an alert on its own: each has an innocent twin, and the value of the signal is that it forces an explanation, not that it proves theft. ## Why every one of them needs a recorded shape 1. There is no absolute threshold. "Forty thousand reads" is alarming for a worker that needs two and unremarkable for a gateway that fetches per request. 2. The comparison is always **per identity**. An estate-wide average hides exactly the worker you are looking for, because one identity's departure disappears into the fleet's spread. 3. The shape has to be recorded before the incident. Reconstructing "what was normal" after the fact, from the same record the attacker has been writing to, is the position nobody wants to be in. This assumes the reads are recorded at all, with the caller, the source, the time, the name and the outcome on each event. Whether the store produces that record, and what it takes to keep it trustworthy, is a separate problem from reading signals out of it. ## What the signal is, and is not A departure is a reason to ask a question, and the honest posture is that it is evidence, not a verdict. The engineering claim worth making in an interview is narrower and stronger than "monitor for anomalies": **the class of abuse you are defending against consists entirely of permitted reads, so any control that only examines refusals has zero coverage of it.** Once you say that, the follow-on design work — which dimensions you record per identity, and what you do when one of them moves — has an obvious shape.

  • If denial events miss the abuse, is there any reason to keep alerting on them?
    Yes, for a different class of problem. Refusals show a caller reaching for names it was never granted, a consumer still presenting a withdrawn value, or a credential aimed at the wrong branch of the name space. They are cheap and specific. They simply have no coverage of a credential the store honours.
  • What makes one successful read distinguishable from another at all?
    Only its context relative to that identity's established reads: which names, how many, how often, from where, at what hour. Nothing inside a single event separates the thief's read from the owner's, because both are the same operation with the same credential. The comparison is what creates the signal.
  • The attacker reads the value once, during business hours, from inside the fleet's own address space. What fires?
    Probably nothing, and that is the honest answer. A single in-shape read is indistinguishable from normal work. These signals bound how long undetected abuse can continue and how much of it can happen quietly; they do not promise that the first use is caught.

A stolen building pass opens the door silently, exactly like the real one, so the reader at the door has nothing to report. What gives it away is the pass appearing in a wing its holder has never worked in, at three in the morning, twenty times in a night.

saying these in an interview costs you the question

  • Claims a failed-login alert would have caught the stolen credential
  • Treats a successful read as proof that it was legitimate
  • Says watching denial events alone is sufficient monitoring
  • Proposes rate-limiting authentication attempts as the fix
  • Assumes the store can tell a thief from the credential's owner
  • Says 'monitor for anomalies' without naming a single dimension