skip to content

The platform's own identity service is down, so credential refresh fails — why do your services keep serving at first, and what ends that?

level: middleimportance: must knowfreq 62%

answer

  1. the refresh path, not the call path
  2. already-issued credentials are still valid
  3. grace runs out at expiry
  4. a cliff, not a ramp
  5. synchronized issue times, synchronized cliff

basics

~20 s

Only the refresh path is broken. A workload still holding an unexpired short-lived credential keeps calling platform services normally. Serving stops at the expiry cliff, when that credential's remaining lifetime runs out — often for many workloads within the same few minutes.

solid answer

~50 s

A short-lived platform credential has an issue time and an expiry, and it is fetched from the provider's own identity service and refreshed on a schedule well before it runs out. When that identity service is unavailable, the refresh call fails but the credential a process already holds is still valid, and on most platforms the services accepting it validate it without a live call back to the issuer. So the first symptom is a log line, not an outage. The real failure arrives at the **expiry cliff**: the moment the held credential expires, every platform call from that process is rejected. It is a step, not a ramp — and because a fleet deployed in one window refreshed in one window, the steps line up. Anything that starts up or restarts during the outage fails immediately, since it has nothing to fall back on.

code

pseudocode · 28 lines
pseudocode
credential = null
refreshInFlight = false

function currentCredential():
    now = clock.now()
    if credential != null and now < credential.expiresAt:
        if now >= credential.refreshAfter and not refreshInFlight:
            startBackgroundRefresh()      // a failure here is not fatal yet
        return credential                 // still valid: keep serving
    // held credential is missing or expired - the cliff
    credential = identityService.fetch()  // only path that can fail hard
    scheduleNextRefresh(credential)
    return credential

function startBackgroundRefresh():
    refreshInFlight = true
    try:
        fresh = identityService.fetch()
        scheduleNextRefresh(fresh)
        credential = fresh
    catch error:
        log("refresh failed; serving on held credential until " + credential.expiresAt)
    finally:
        refreshInFlight = false

function scheduleNextRefresh(c):
    halfway = c.issuedAt + 0.5 * c.lifetime
    c.refreshAfter = halfway + random(0, 0.1 * c.lifetime)   // jitter spreads the fleet

go deeper

for a junior

Know that a workload does not hold a permanent password to the platform: it gets a short-lived credential that expires and is renewed automatically. If renewal fails, what it already holds still works until it expires.

for a middle

Explain the mechanics: issue time, lifetime, refresh scheduled well before expiry, and local validation by the accepting service. Be able to say why the failure is a step at expiry rather than a gradual rise in errors.

for a senior

Show you would alarm on time-until-expiry rather than on refresh errors, add jitter so the fleet does not fall together, and avoid restarting processes during the window because a restart has nothing held to fall back on.

for a principal

Frame the trade you are actually making: a short lifetime bounds the value of a stolen credential and buys a correlated cliff; a long one does the reverse. Decide the estate's lifetime deliberately and say which product paths must survive it.

## The dependency nobody draws Every workload on a large cloud platform proves what it is to that platform before it can do anything useful — read from an object store, write to a managed database, publish a message. On a modern estate it does that with a **short-lived credential**: an artifact carrying an issue time and an expiry, obtained from the provider's own identity service and refreshed on a schedule long before it lapses. The refresh is invisible. It has no dashboard, no owner on the team, and it usually lives inside a platform client library that application code never calls directly. That makes the platform's identity service a **shared dependency**: one service, operated by the provider, that every tenant's every workload touches continuously. When it fails it does not fail for you — it fails for everyone at once. That is the defining property of this failure class and it is why the reflexes learned on ordinary outages do not apply. There is nowhere to shift traffic that is not also calling the same service, and the usual mitigation of replacing unhealthy processes makes things strictly worse. ## Why the first minutes look normal A short-lived credential is a **bearer artifact**. On most platforms the service that accepts it validates it locally — signature and expiry — rather than asking the issuer whether it is still good. Platforms differ in how much of that validation is local and how much is live, but the practical shape is the same everywhere: the refresh path breaks before the call path does. So in the opening minutes: - a process holding an unexpired credential keeps calling platform services and keeps succeeding; - a process whose scheduled refresh fails writes an error and carries on with what it already has; - a process that is starting up, or that has never held a credential, fails at once — it has nothing to fall back on; - your customers see nothing, and your error dashboards are flat. ## The cliff What follows is not degradation. A credential is fully valid right up to its expiry and worth nothing one instant later, so the failure for any given process is a step from roughly zero errors to total rejection. | Moment | What the workload sees | What customers see | |---|---|---| | Identity service fails | Refresh calls error; the held credential is still valid | Nothing | | Between failure and expiry | Repeated refresh failures in logs; a process that restarts cannot come back | Nothing, unless a process restarts | | At expiry | Every platform call rejected as unauthenticated | Total failure for that process | | After the service returns | Refresh succeeds, possibly only after a retry storm subsides | Recovery, sometimes slower than the outage itself | The last row is the one teams are least ready for. When the identity service comes back, every workload that has been retrying is waiting on it simultaneously, and an unbounded retry loop turns a recovering service into an overloaded one. ## Why a whole fleet goes together The cliff is per-process, but the processes rarely fall independently. Three things line them up: 1. **Shared deployment time.** A fleet rolled out in one window fetched its first credential in one window, so every subsequent refresh is on the same clock. 2. **A single lifetime policy.** The credential lifetime is usually one value applied across the estate, so every process has the same amount of grace. 3. **No jitter.** Refresh schedules derived deterministically from the issue time reproduce the alignment on every cycle. The result is correlated failure inside your tenant stacked on top of correlated failure across tenants — a narrow window in which almost everything expires. ## What you can actually do - **Refresh ahead, with jitter.** Refresh at a fraction of the lifetime rather than near its end, and add a random offset so the fleet's expiries spread out instead of stacking. - **Treat a failed refresh as a countdown, not an error.** The useful alarm is *time until the oldest held credential expires*, which tells you how long you have; a refresh-error rate tells you only that something started. - **Retry with backoff and jitter.** This protects the identity service on the way back up, which is when it is most fragile. - **Know which product paths need a platform call at all.** Paths served entirely from memory or from an already-open connection survive the cliff; naming them in advance is what makes a deliberate degraded mode possible instead of an improvised one. - **Do not conclude that long-lived static keys are the fix.** They genuinely would have survived this outage. They buy that by staying valuable to an attacker forever, which is a far worse trade than a rare correlated failure of the provider's identity service.

  • What makes an entire fleet expire within the same few minutes, and how would you spread it out?
    Processes deployed in one window fetch their first credential in one window, and a lifetime policy applied estate-wide gives them all the same grace. A deterministic refresh schedule then reproduces the alignment forever. Add a random offset to the refresh time so each process drifts onto its own clock, and the fleet's expiries spread across the window instead of stacking into one.
  • Once the identity service returns, why can recovery take longer than the outage did?
    Every workload that has been retrying is waiting on the same service, so it comes back into a thundering herd. If clients retry in tight loops, the recovering service is immediately overloaded and flaps. Exponential backoff with jitter on the refresh path is what turns that into a gradual ramp, and it is worth more than any amount of extra retry aggression.
  • What would you actually have to change to keep serving past the cliff?
    Either lengthen the credential lifetime, which directly lengthens how long a stolen credential is useful, or remove the platform call from the path. The second is the real answer: identify which product paths need a platform service on the request path at all, and make the rest — cached reads, work already queued, responses served from memory — able to complete without one.

A building keycard system whose card printer and door readers are separate: when the printer is down, everyone already carrying a card keeps opening doors until their card expires at midnight, and nobody can get a new one. New starters are locked out immediately.

saying these in an interview costs you the question

  • Assumes everything stops the instant the platform's identity service fails
  • Concludes long-lived static keys are the safer design because they never refresh
  • Calls it a total outage while the data path is still serving normally
  • Expects retrying the refresh harder to eventually produce a credential
  • Confuses the identity service being unavailable with permissions being revoked
  • Ignores that a restarting process has nothing held to fall back on