Your platform must publish one default staleness bound for every service that caches a fetched credential — what evidence sets that number?
answer
- the promise sets the ceiling
- fleet size over the bound sets the floor
- scale it with what the credential reaches
- largest bound that keeps the promise
- deviation is a request, not a constant
basics
~20 sTwo measured quantities: how fast the organisation has promised it can make a credential stop working, and the store read rate the fleet size divided by the bound produces. The bound is the largest number that keeps both promises.
solid answer
~50 sStart from the commitment, not from taste. If the incident promise is "a withdrawn credential stops being used within twenty minutes", the bound is the dominant term in that number and everything else — refresh attempts, in-flight calls — is slack you must subtract. Then check the arithmetic in the other direction: process count divided by the bound is the read rate every service will put on the store, so a fleet of thousands at a one-minute bound is a load decision, not a security one. Two more inputs adjust it: what the credential can reach if it is used after withdrawal, and whether a faster signal already exists, because a rejection-triggered refetch converges an active holder in one call and makes the bound a backstop rather than the primary control. Publish one default so the promise is computable, and make deviation a request with evidence.
go deeper
Know that the interval between refreshes of a held credential is a deliberate number someone chose, not a framework default, and that it decides how long a stale value keeps being used.
Do both arithmetics: the withdrawal promise minus refresh and in-flight time gives the ceiling, and fleet size divided by the bound gives the store read rate it costs. Be able to state both for a concrete fleet.
Argue the number for a real service — blast radius of the credential, whether a rejection trigger already converges active holders, what a refresh costs the request path — and show how you would verify the bound is actually held rather than merely configured.
Own it as an estate-wide commitment: one published default makes the incident question answerable, deviation becomes a reviewed request with evidence, and the effort goes into eliminating holders with no trigger, since they are what really sets the window.
A staleness bound is the one number in a cached-credential design that somebody has to argue for. It is tempting to pick it by feel — five minutes sounds careful, an hour sounds reasonable — and the result is an estate where nobody can answer how quickly a credential can be made to stop working. Publishing one default is what makes that question answerable. ## The two arithmetic constraints Everything else is judgment; these two are computation. 1. **The withdrawal promise sets the ceiling.** If the organisation's incident commitment is that a withdrawn credential stops being used within twenty minutes, then the bound plus the worst refresh attempt plus the longest in-flight call must fit inside twenty minutes. With a one-minute refresh cost and two-minute calls, the bound can be at most seventeen — not twenty. 2. **The fleet size sets the floor.** Read rate is roughly *processes ÷ bound*. Six hundred processes on a sixty-second bound is ten reads a second for one credential; drop the bound to ten seconds and it is sixty. Multiply by the number of distinct credentials and the number of services adopting the default, and the floor is whatever the store can serve without the refreshes themselves becoming the availability problem. The defensible default is the **largest** bound that satisfies the first constraint, because every unit below it is store load bought for nothing. ## The judgment inputs - **What the credential reaches.** A read-only key for a scoring service and a credential that can write to customer records deserve different windows. The bound is a blast-radius control, so it should scale with blast radius. - **Whether the credential expires on its own.** If the store mints short-lived per-consumer credentials, their own validity is already a ceiling, and a separate bound below it adds reads without shortening anything. The default should say so explicitly rather than being applied blindly on top. - **Whether a faster signal exists.** Where holders refetch on rejection, active holders converge in one call and the bound is the backstop for idle ones. Where no such signal exists, the bound is the only control and has to carry the whole promise. - **What a refresh costs the service.** A refresh on the request path of a latency-sensitive service is a different cost from a background one. - **How many holders there are, and whether they all have triggers.** A default bound applied to a fleet where some processes resolve the value once at start-up is a promise the platform cannot keep, and finding those holders is worth more than any choice of number. ## Why one default rather than per-service choices | one published default | every service picks its own | |---|---| | the estate's window is one number you can quote | the window is the worst service, and nobody knows which | | store load is predictable from fleet size | load is discovered during an incident | | a deviation is visible and reviewable | every value is equally unexplained | | new services inherit a defensible choice | new services inherit whatever was copied | The right of a team to deviate should exist, but as a **request with evidence**, not a silent constant: 1. **Longer than the default** is granted where a compensating trigger exists — rejection-driven refresh in a service that calls constantly — or where the credential's own expiry is already shorter than the default. 2. **Shorter than the default** is granted where the blast radius justifies it *and* the read-rate arithmetic has been done for that fleet size. 3. **No bound at all** is never granted, because that is the unbounded holder, and it is the case that makes the estate-wide promise false. ## What the default is not It is not a rotation schedule: rotation decides how long a value stays in use, while the bound decides how fast a replacement or a withdrawal is felt. It is not a substitute for withdrawal either — letting copies age out is not the same act as making the other side refuse the value, and a plan that relies on the bound alone has chosen to wait rather than to act. And it is not a security control on its own: it bounds the *time* a stale value is used, not *what* that value can do. ## The principal's version of the answer The value of publishing the number is not the number. It is that during an incident someone can say "eighteen minutes, and here is the fetch record that shows the oldest holder" instead of "we think most services refresh fairly often". Set the default from the promise, check it against the read rate, scale it with blast radius, make deviation a reviewed request, and spend the effort you save on finding the holders that have no trigger at all — they are what actually decides the estate's window.
- When may a team hold a credential for longer than the published default?When something else carries the promise. A service that calls constantly and refetches on rejection converges in one call, so its effective window is not the bound. So does a credential whose own expiry is shorter than the default. Both are evidence a reviewer can check; "our service is special" is not, and no bound at all is never granted.
- A team wants a five-second bound for a fleet of two thousand processes. What do you ask them?For the read-rate arithmetic first: two thousand divided by five is four hundred reads a second for one credential, before any other service adopts the same reasoning. Then ask what the five seconds is buying that a rejection trigger would not, since that converges an active holder in one call at no steady-state cost.
- How do you tell whether the published default is actually being kept?Measure it from the store's fetch records rather than from the configuration: the oldest last-fetch time per holder is the observed window, and a holder that appears once and never again is the unbounded case. A default that is configured everywhere and observed nowhere is documentation, not a control.
saying these in an interview costs you the question
- Picks the bound by feel with no arithmetic behind it
- Sets it far below what the store's read budget can serve
- Treats the bound as the rotation schedule
- Relies on ageing copies out instead of withdrawing the value
- Applies one number without regard to what the credential reaches
- Lets each service choose silently, so the estate's window is unknown