skip to content

Half your users' stored grocery-list grants expire within the same hour — do you refresh them on a background sweep or lazily at first use?

level: seniorimportance: should knowfreq 44%

answer

  1. who is waiting for the token
  2. dormant users cost calls, not latency
  3. expiry arrives as a cohort
  4. jitter the due time, not the expiry
  5. one row, one refresh in flight

basics

~20 s

Refresh at the moment of use for anything a user is waiting on, and sweep ahead only for the grants background work needs; drive the sweep off a jittered due time so a cohort authorized in one hour cannot stampede the provider.

solid answer

~50 s

Both, on different paths. Lazy refresh wastes nothing on dormant users but pays a provider round trip inside the request that needed the grant, and it discovers a dead grant in front of a user. A sweep keeps tokens warm for work that runs with nobody present, but refreshes people who left months ago and concentrates load wherever the schedule points. The cohort is the real problem: users authorized in the same hour receive the same `expires_in` and stay synchronised, because every refresh re-synchronises them. So store a `refresh_due_at` at expiry minus a random slice of the lifetime, re-randomise it on every refresh, and let the sweep filter on that column with a batch limit. Hold a short per-row lock — a provider that rotates the refresh token turns two concurrent refreshes into one destroyed grant — and defer the whole batch, honouring `Retry-After`, when the provider rate-limits you.

code

pseudocode · 26 lines
pseudocode
# every 5 minutes; dead rows are never selected
due = grants.where(status == ACTIVE and refresh_due_at <= now).limit(BATCH)

for grant in due:
    if not try_lock(grant.id, ttl = 60s):
        continue                      # a user-facing refresh already has this row
    try:
        r = provider.refresh(decrypt(grant.refresh_token))
    except RateLimited as e:
        defer(due, by = e.retry_after + random(0s, 60s))
        break                         # back off the whole pass, not one row
    except InvalidGrant:
        grant.status = DEAD           # terminal: no retry, no backoff
        grant.status_reason = "refresh refused: invalid_grant"
        grant.access_token = null; grant.refresh_token = null
        grant.save(); continue
    except Timeout or ServerError:
        grant.refresh_due_at = now + random(1min, 5min)
        grant.save(); continue        # transient: try again soon

    grant.access_token = encrypt(r.access_token)
    if r.refresh_token != null:       # provider rotated it
        grant.refresh_token = encrypt(r.refresh_token)
    grant.access_token_expires_at = now + r.expires_in
    grant.refresh_due_at = grant.access_token_expires_at - random(10min, 30min)
    grant.save()                      # one transaction with the response

go deeper

for a junior

Know the two options exist — refresh when something needs the grant, or refresh ahead of time on a schedule — and that the stored expiry is what either one consults.

for a middle

Explain who pays in each case: the caller waits in the lazy path, while a sweep spends provider calls on users who may never return.

for a senior

Bring up the cohort. Show the jittered due time, the batch limit, batch-level backoff on a rate-limit response, and the per-row lock that protects a rotated refresh token.

for a principal

Frame it as a capacity question against somebody else's rate limit: which populations must be warm at all, what the integration is allowed to spend per hour, and when a dormant grant should be dropped rather than maintained.

## The two schedules, stated honestly **Lazy** means: when something needs the grant, read the stored expiry, refresh if it has passed or is close, then do the work. **Sweep** means: a job walks the table on a schedule, finds rows approaching expiry, and refreshes them before anyone asks. They are not a style preference. They fail differently, and they fail at different people. | | lazy, at first use | background sweep | |---|---|---| | who waits for the provider | the caller — a user, or the job that needed it | nobody | | dormant users | cost nothing | refreshed forever, for no one | | load shape | follows your traffic and peaks with it | whatever the schedule says | | provider slowness | surfaces as a slow or failed user action | surfaces as a queue you can drain later | | where a dead grant is discovered | in front of a user | in the sweep, before anyone notices | A planner that writes shopping lists overnight cannot be purely lazy: the token has to be valid when the job runs and no user is there to absorb a refresh. A planner whose only third-party calls happen while someone taps a button barely needs a sweep at all. ## Why a whole population expires together Expiry is not naturally spread out. It is minted in cohorts: - a launch, or the day a feature that needs the grant becomes visible; - a campaign that drives thousands of people through authorization in an afternoon; - a mass re-authorization after your client credential changed; - a bulk reconnect after an incident. Every member of a cohort receives the same `expires_in`, so their access tokens expire within minutes of each other — and stay that way, because a refresh resets every clock in the cohort to the same new value. The cohort does not decay; it re-forms on every cycle. That is the shape that turns a sweep into an incident: ten thousand refreshes arriving inside one minute, against an authorization server whose rate limit, not your worker count, is the binding constraint. ## Jitter, and what the due time actually is 1. Store a **`refresh_due_at`** alongside the absolute expiry, set to expiry minus a random slice of the token's lifetime — a spread of tens of minutes on an hour-long token. 2. **Re-randomise it on every refresh.** A fixed offset moves the stampede; a fresh random offset dissolves it over a few cycles. 3. **The sweep filters on `refresh_due_at`**, never on raw expiry, and always with a batch limit so one pass cannot become an unbounded burst. 4. **Treat a rate-limit response as a batch-level signal.** Defer the remaining rows by `Retry-After` plus jitter and stop the pass; retrying row by row is how a throttle becomes an outage. ## One row, one refresh in flight Providers vary: some hand back a new refresh token on every use, some do not. Where one does, it usually supersedes the previous refresh token at the moment it issues the replacement, and that makes concurrency sharp: - two workers, or a worker and a user-facing request, send the same refresh token at once; - one response arrives with a replacement, the other is refused; - whichever result is persisted last wins, and if that is the loser's, the row now holds a token the provider has already superseded. The grant is then dead, and the user has to authorize again — an outage you inflicted on yourself. Take a short per-row lock before refreshing, persist the new access token, the new refresh token and the new expiry in one transaction with the response, and skip the row rather than queue behind the lock. ## What the sweep must never do - **Refresh rows already marked dead.** They cost a request each and the answer never changes. - **Retry a refusal that means the grant is gone.** `invalid_grant` is terminal; backoff is for timeouts and `5xx`. - **Keep a dormant user warm indefinitely.** Someone who has not opened the planner in six months needs no fresh token; refresh on demand and let a retention rule clear the row. - **Scale its way out of a rate limit.** More workers against a throttling provider produces the same throughput and more errors. ## The answer to give Say which path each strategy serves — lazy for user-present work, sweep for the grants background jobs need — then spend the rest of the answer on the cohort: jittered due time, batch limits, batch-level backoff, one refresh in flight per row, and dead rows excluded. That sequence is what separates someone who has run an integration from someone who has read about one.

  • Why is raw expiry the wrong column to drive the sweep off?
    Because expiry is synchronised by construction. Users authorized in the same hour get the same `expires_in`, and every refresh re-synchronises them, so selecting on expiry reproduces the cohort on every cycle. A separate due time, randomised inside the token's lifetime and re-randomised on each refresh, spreads the same work across the hour.
  • A worker refreshed a rotating grant and crashed before persisting the response — what state is the row in?
    Probably dead. The provider issued a replacement and, at a provider that rotates, usually invalidated the stored one as it did so, so the next refresh is refused with `invalid_grant`. That is why the persist belongs in the same transaction as the response and why one row holds one refresh at a time; recovery is putting the user through authorization again.
  • What do you do about a user who has not opened the planner in six months?
    Stop refreshing them. Sweeping a dormant grant burns rate limit you need for active users and keeps a live credential warm for nobody. Refresh that row lazily if they come back, and let a retention rule clear grants unused past a stated window — the row is standing access to someone's account, so keeping it warm is a cost, not a courtesy.

saying these in an interview costs you the question

  • Sweeps every stored grant every few minutes regardless of use
  • Schedules the refresh exactly at expiry, so a cohort refreshes in one minute
  • Retries a grant the provider already refused, on timeout-style backoff
  • Lets two workers refresh one row at once where the provider rotates
  • Treats the provider's rate limit as something more workers can fix
  • Assumes a lazy refresh is free because nobody scheduled it