skip to content

A scoring credential is withdrawn at the provider while forty workers each hold a cached copy — how long does it keep working?

level: seniorimportance: must knowfreq 50%

answer

  1. the clock starts at the accepting side
  2. maximum over holders, never the mean
  3. bound plus refresh plus in-flight
  4. a holder with no trigger never converges
  5. deleting the record withdraws nothing

basics

~20 s

Until the slowest holder refreshes, not the average one: roughly the staleness bound plus one refresh attempt plus the longest call already in flight. A holder with no refresh trigger at all keeps working indefinitely.

solid answer

~50 s

The clock starts when the provider stops accepting the value, not when someone edits or deletes the record in the store. From that instant each worker keeps presenting the copy it holds until its own trigger fires. For a worker on a fifteen-minute bound that refreshed one second before the withdrawal, that is up to fifteen minutes, plus whatever a failed refresh costs in retries, plus the longest call that was already in flight. The fleet's window is the **maximum** over holders, so it is set by the worst one: a worker that refetches only on a rejection converges when it next makes a call, which for an idle worker can be hours, and a process that fetched once at start-up and has no trigger at all never converges until it restarts. Quote the maximum, never the mean.

code

pseudocode · 13 lines
pseudocode
# worst case, measured from the moment the provider stops accepting the value
function exposureWindow(holder):
    if holder.refreshTrigger == NONE:
        return UNBOUNDED                 # fetched once at start-up, never again
    if holder.refreshTrigger == ON_REJECTION_ONLY:
        return holder.timeUntilNextCall + holder.callDuration   # idle holders wait
    # bounded refresh: assume it fetched an instant before the withdrawal
    return holder.bound
         + holder.refreshAttemptCost     # includes one backoff if the first fetch fails
         + holder.longestInFlightCall

fleetWindow = UNBOUNDED if any(h.refreshTrigger == NONE for h in holders)
              else max(exposureWindow(h) for h in holders)

go deeper

for a junior

Recall that a cached credential keeps working after it is withdrawn, and that how long depends on when the holder next goes back for a fresh value. The number is not zero.

for a middle

Explain the terms: the bound, the refresh attempt, and the calls already in flight. Be clear that deleting the record in the store is not the same act as making the other side stop accepting the value.

for a senior

Compute it for a real fleet and quote the maximum, then name the holders that break the formula — the idle worker on a rejection-only trigger and the process that resolved the value at start-up. Say how you would evidence the number from fetch records.

for a principal

Treat the window as a commitment the organisation makes and must be able to prove. Decide what the estate's number is, what it costs in store read rate, and whether it is cheaper to buy the guarantee structurally with credentials that expire on their own.

This is the arithmetic question behind every cached credential, and the one an incident commander actually asks: the credential has been made to stop working — for how long is it still working anyway? ## Start the clock in the right place The window opens when **the accepting system stops accepting the value** — the provider, the downstream account, whatever evaluates the credential. It does not open when the record is edited or removed in the secret store. Removing the source stops future fetches from returning it; it does nothing to the copies already delivered and nothing to the other side's willingness to take them. If the provider still accepts the old value, deleting it from the store has shortened no window at all — it has only guaranteed that the next refresh finds nothing. ## The per-holder arithmetic For one holder, the worst case is: **(time until its trigger fires) + (time for the refresh to succeed) + (the longest call already in flight with the old value)** Each term matters: 1. **Time until the trigger fires.** For a bounded refresh this is at most the whole bound, because the unlucky holder refreshed an instant before the withdrawal. Fifteen-minute bound, worst case fifteen minutes. 2. **Time for the refresh to succeed.** A refresh that fails and backs off adds its retry interval. If the fetch takes two attempts thirty seconds apart, add thirty seconds. 3. **The longest in-flight call.** A request that had already picked up the old value and is mid-flight is still exposure. For a long scoring call or a batch submission, this can dominate everything else. So a forty-worker fleet on a fifteen-minute bound with a thirty-second retry and calls up to a minute long is about **sixteen and a half minutes**, not fifteen — and that is the good case. ## The fleet's window is the maximum, not the mean Averaging is the classic mistake. With a uniform bound the *mean* holder converges in half the bound, but nobody is protected by the mean: one worker still presenting a withdrawn credential is still an open window. Worse, real fleets are not uniform: | holder shape | when it converges | window | |---|---|---| | bounded refresh | at its next tick | at most the bound, plus refresh and in-flight | | rejection-triggered only | at its next call that is rejected | unbounded for an idle holder | | fetched once at start-up | at its next restart | as long as the process lives | | a copy handed to a third party | when that party is told | outside your control entirely | The last two rows are why the honest answer to "how long?" starts with "which holders exist?" A single long-running process that resolved the value at start-up and never again turns a fifteen-minute promise into a several-week reality, and it will not show up in any dashboard of refresh rates because it makes no refreshes to count. ## What shortens the window - **Lower the bound.** Directly proportional, and directly proportional in store read rate too — halving the bound doubles the reads. - **Give every holder a trigger.** The unbounded rows above are worth more than any amount of tuning on the bounded ones. - **Make rejection a trigger as well as the bound.** It converges an active holder in one call. - **Prefer credentials that expire on their own.** A short-lived, per-consumer credential bounds the window by construction: nobody has to remember to refresh, because the value stops working on its own. - **Reduce the number of holders.** Fewer copies means fewer independent clocks and a smaller maximum. ## What does not shorten it - **Deleting the value in the store.** It is not withdrawal. The other side decides. - **Writing a new value into the store.** That is replacement; unless the old value was also withdrawn, both may work. - **Rotating on a schedule.** A schedule bounds how long a value is *in use*, which is a different promise from how fast a compromised one can be *stopped*. ## Saying it in an interview Give the formula, then the number, then the caveat: "worst case is the bound plus a refresh plus the longest in-flight call, so about sixteen minutes for this fleet — unless a holder has no refresh trigger, in which case there is no bound at all and that is the first thing I would go and check." That sequence — arithmetic, number, the holder that breaks the arithmetic — is what separates someone who has run this from someone who has read about it.

  • What shortens that window, in the order you would actually do it?
    First find and fix holders with no trigger, because they are unbounded and no tuning helps them. Then add rejection as a second trigger, which converges active holders in one call. Then lower the bound, paying for it in store read rate. Structurally, move to credentials that expire on their own, so the window is bounded by the credential rather than by everyone's diligence.
  • The team deleted the value in the secret store as soon as the incident opened. What did that buy?
    Little, and it can cost. Deletion stops future fetches from returning the value; it does not make the provider refuse copies already held, so the window is unchanged. Meanwhile refreshes now fail, which on some holders means they keep serving the old copy rather than replacing it — so the act aimed at shortening the window can extend it.
  • Does knowing the average refresh age of the fleet tell you the window?
    No. The window is the maximum over holders, and the mean hides exactly the holders that set it — the idle worker and the process that never refreshes. A mean refresh age of seven minutes is compatible with one worker that has held the same value for three weeks.
  • How would you evidence the window rather than assert it?
    From the store's record of fetches per holder: the oldest last-fetch time across the fleet is the observed floor for the window, and any holder that appears once and never again is the unbounded case. Pair it with the longest observed call duration for the in-flight term. That turns a design claim into a measured number you can quote during an incident.

saying these in an interview costs you the question

  • Quotes the mean refresh age as the exposure window
  • Starts the clock when the value was deleted in the store
  • Assumes every holder refreshes on the same schedule
  • Forgets calls already in flight with the old value
  • Thinks writing a new value into the store stops the old one
  • Never asks whether some holder has no refresh trigger at all