Who holds the renewal clock for a leased credential, and how early should it fire relative to the expiry?
answer
- one owner, never two
- fire on a fraction, not the deadline
- what remains is a retry budget
- aligned fleets renew in one second
- permanent refusal is not a retry
basics
~20 sExactly one owner: the process holding the credential, or the companion beside it that fetched it — never both. It fires partway through the window, early enough that several failed attempts and a store blip still fit before the expiry.
solid answer
~40 sPick one owner per lease. Either the process that uses the credential renews it, or a companion beside it does and the process only re-reads what it is given; two owners renewing the same lease produces double the load and no clear answer to "who stopped renewing?". Fire the clock at a **fraction** of the window — around two-thirds through is a common shape — so that the remaining third is a retry budget covering a few attempts with backoff and a short store outage. Renewing at the expiry leaves no budget at all. Across a fleet that redeploys together, add a little spread to the first renewal, or every consumer that started in the same minute will renew in the same second forever.
code
pseudocode · 19 lineswindow = lease.expiresAt - lease.grantedAt // use what was GRANTED
renewAt = lease.grantedAt + 0.66 * window + jitter(0, 0.05 * window)
safety = 0.1 * window
wait until renewAt
attempt = 0
while now < lease.expiresAt - safety:
result = store.extend(lease.handle)
if result.granted:
lease = result.lease // re-derive the next fire time
scheduleRenewal(lease)
return
if result.refusedAtCeiling:
break // no retry can succeed
attempt = attempt + 1
wait backoff(attempt)
credential = store.issue(consumerRole) // different value
reconnectDownstream(credential)go deeper
The idea to take away is that something must actively extend a leased credential before its deadline, and that waiting until the deadline leaves no room for the attempt to fail.
Explain the fraction-of-the-window rule and why what remains is a retry budget, and know that the schedule must come from the deadline the store actually granted rather than the one requested.
Demonstrate the operational shape: one owner per lease, spread across a fleet that deploys together, a distinct branch for a permanent refusal, and the per-lease signals that reveal a holder which quietly stopped renewing.
The judgment is whether the clock belongs in every consumer or in one companion the platform owns. Centralising makes it correct once and monitorable, at the price of a new dependency whose silent death is invisible to the workloads it serves.
A lease is a promise with a deadline, and someone has to act before that deadline. "The renewal clock" is that piece of the contract: who owns it, when it fires, and what it does when the store says no. ## Exactly one owner per lease The renewal call can sit in two plausible places, and the failure mode is having both: - **In the consumer itself.** The process that holds the credential also extends it. Simple, no extra moving part, and the process knows whether it still needs the credential at all. The cost is that every consumer, in every codebase, carries timer and retry logic, and each gets it slightly wrong. - **In a companion process beside the workload.** One implementation of the clock serves many consumers, which is why the pattern exists. The cost is a second thing that can die quietly: if the companion stops, the consumer keeps working perfectly right up to the expiry it never heard about. What must not happen is both, or neither. Both means two renewal streams against one lease — twice the load on the store, and no meaningful answer to "has anything renewed this lease recently?", which is the signal that tells you a holder has gone. Neither happens more often than anyone admits: a credential fetched once at start-up by a helper that was never asked to keep it alive. ## How early the clock fires Renewal at the moment of expiry is not renewal; it is a coin toss. Fire at a fraction of the window so that what remains is a budget: - **Around two-thirds through the window** is a common shape: it leaves roughly a third of the lifetime for failures, and it keeps the call rate modest. - The remaining time must cover several attempts with backoff **plus** a plausible short outage of the store, not one attempt. - Compute the schedule from the window the store actually granted, not from what was requested — a store may grant less than you asked for, and a clock built on the request will fire late. - Re-derive the next fire time from the **new** expiry after each success, rather than adding a fixed interval to the old one, so drift does not accumulate. ## Spread the fleet Consumers that start together renew together. A fleet redeployed in one rollout has every holder's clock aligned to the same second, and that alignment persists for as long as the processes live, because every renewal returns the same window length. Add a small random offset to the **first** renewal — a few percent of the window is enough — so the fleet's renewals arrive spread out. This is worth doing even when the store would cope with the burst, because a synchronised fleet also means every consumer fails at once during a short store outage, turning a blip into a correlated event. ## What the clock does when the answer is no Two refusals look alike at a glance and must be handled differently: 1. **A transient failure** — the store is briefly unreachable or busy. Back off and try again, still inside the budget the clock reserved. 2. **A refusal to extend any further** — the credential has reached the ceiling on its total life. No retry will ever succeed, and the correct response is to obtain a new credential and adopt it. Routing the second case into the first is the classic way a well-instrumented consumer still runs out of validity: it retries a permanent answer until there is no time left. ## What to record The clock is cheap to observe if you decide what it should say: - Outcome of the last renewal per lease, and when it happened. - Time remaining on the current window at the moment of each attempt — a number that should be roughly constant and is a good alarm when it starts shrinking. - Leases with no renewal in longer than one window, which are the ones whose holder has probably gone. ## Where designs differ Some stores tell a holder exactly how much life a renewal granted and how much of the ceiling remains, which makes the clock trivially correct. Others return only a new deadline, and the holder has to track the total age itself. Some cap the length any single renewal may add, so the clock has to fire more often than the original window would suggest. Build the clock against the numbers the store returns, and treat everything else as an assumption to verify.
- Why derive the next renewal from the granted window rather than the one you asked for?Because a store may grant less than requested — a class-wide cap, a policy limit, or the remaining life to the ceiling can all shorten it. A clock built on the request then fires after the credential has already lapsed, and the symptom is a consumer that renews perfectly on paper and still runs out. Read the deadline the store returned and schedule from that.
- A companion process renews on the consumer's behalf. What new failure does that introduce?A silent one. The consumer keeps working on a valid credential while the companion is dead, and nothing is wrong until the window ends. It has to be monitored as its own thing — last successful renewal per lease, and an alarm when a lease's remaining life trends downward instead of staying roughly flat — because the consumer will not report the problem.
saying these in an interview costs you the question
- Renew right before the expiry to get the longest life.
- Both the process and its helper can renew; duplicates are harmless.
- Schedule the next renewal by adding a fixed interval each time.
- Retry a refused extension until it eventually succeeds.
- Jitter is unnecessary because the store can handle the load.