A six-hour data pipeline run started failing halfway through once session lifetimes were shortened — what is the job doing wrong?
answer
- resolved once, held forever
- a credential is a cache with an expiry
- refresh before the expiry, not after
- refused calls here are not a permission problem
- lengthening the lifetime is the wrong fix
basics
~20 sThe job resolves its short-lived credential once at start-up and holds those values for the whole process. When the session expires mid-run, every later call is refused for authentication, not permissions. Re-resolve before the expiry the source returned.
solid answer
~50 sShort-lived credentials come back with an **expiry**, and code that keeps the values but discards the expiry is holding a cache it never invalidates. Halfway through the run the session ends, and from that instant every call is refused — including calls identical to ones that succeeded minutes earlier. That shape is the diagnosis: an expiry hits all calls at a point in time, while a missing permission fails one action from the very first attempt. The fix is to resolve through a function rather than into a variable, keep the returned `expiresAt` next to the values, refresh on a margin before it, and treat a refused credential as retryable — discard, re-resolve, retry. The tempting fix, raising the maximum lifetime until the longest job fits, sets the estate's exposure floor by its worst-behaved consumer.
code
pseudocode · 17 linescached = none
function currentCredential():
if cached is none or now() >= cached.expiresAt - refreshMargin:
cached = credentialSource.resolve() # returns values + expiresAt
return cached
for each unit in workUnits:
attempt = 1
loop:
result = platform.call(unit, currentCredential())
if result is credentialRejected and attempt <= maxAttempts:
cached = none # discard, so the next resolve is fresh
attempt = attempt + 1
continue
break
checkpoint(unit.id)go deeper
Recall that a short-lived credential comes back with an expiry time, and that code holding the values after that moment will simply be refused.
Explain the mechanics: resolve through a function, keep the returned expiry, refresh on a margin ahead of it, and retry once on a refused credential after discarding the cached value.
Diagnose from the shape — everything failing from a point in time, including harmless reads — and reject the one-line fix of raising the lifetime, naming what it costs every other caller.
The trade to own is where renewal complexity lives. Pushing it into clients keeps the estate's exposure short; pushing it into a longer ceiling buys quiet at a cost every caller pays.
## The shape of the failure Start from the symptom, because the symptom is nearly diagnostic. Calls succeed for a while, then **every** call is refused — including calls identical to ones that worked minutes earlier, including harmless reads. Nothing changed in the code path and nothing changed in the grant. A failure that starts at a point in time and affects everything at once is an **authentication** failure: the platform no longer accepts the credential being presented. A failure of **authorization** looks different — it affects one action, and it is there from the first attempt. | expired credential | missing permission | |---|---| | begins partway through the run | present from the first attempt | | affects every call, including previously successful ones | affects the one action that is not granted | | harmless read-only calls fail too | harmless calls keep succeeding | | fixed by re-resolving the credential | fixed by changing the grant | The second column is where teams go wrong under time pressure: the run is broken, someone widens the grant, nothing improves, and the estate is left with a permission it did not need. ## Why the job snapshots its credential Resolving a short-lived credential returns two things: the values, and the moment they stop being valid. Code that keeps only the first has silently made a decision — that the credential is a constant. That decision is invisible while lifetimes are long, and becomes a defect the day somebody shortens them. The usual forms are: - reading the credential once at start-up and passing the strings down through the call stack; - capturing the values into process state that is set up once; - building a client object at start-up with the values frozen in, then reusing it for the whole run. None of these is wrong in a five-minute job. All of them break in a six-hour one. ## Treat it as a cache with an expiry 1. **Resolve through a function**, never into a variable read once. The call site asks for a credential every time and does not know whether it was cached. 2. **Keep the expiry the source returned** alongside the values. The lifetime is data from the issuer, not a constant you may assume. 3. **Refresh on a margin**, not at the expiry. A call already in flight, a slow issuer and clock skew between your host and the platform each eat into the tail. 4. **Treat a refused credential as retryable**: discard the cached value, re-resolve, retry the call once, and only then fail. That covers the case where the credential became invalid earlier than its stated expiry. Most platform client libraries already do all four when you let them resolve the credential themselves. The bug is nearly always code that took the values out of the library's hands and carried them itself. ## The fix that makes it worse Raising the maximum session lifetime until the longest job fits is the tempting one-line change, and it inverts the design: the estate's exposure window is now set by its worst-behaved consumer, and every other caller — including interactive ones on laptops — inherits it. It also does not scale, because the next job is longer. ## The ceiling a refresh cannot cross Platforms commonly cap the **total** life of a session independently of how long any one issued credential is valid, and designs differ in how the two interact. Where such a cap is shorter than the unit of work, refreshing cannot save the job, and the honest answers are structural: - make the unit of work smaller, with a checkpoint after each unit and a resume that skips finished ones; - have each unit run under a freshly acquired session rather than stretching one across the whole run; - treat the caller that genuinely cannot do either as a named exception with an owner, not as a reason to raise the ceiling for everybody. ## What good looks like - The credential is resolved lazily and logged by **expiry time**, never by value. - A refresh failure is a distinct, alertable event, separate from the platform call failing. - Long runs checkpoint, so an expiry costs a retry rather than six hours of recomputation. - Failure messages distinguish *refused because the credential expired* from *refused because the identity may not do this*, so the next person diagnoses in minutes and does not widen a grant to fix a clock. The underlying point of the leaf holds throughout: a lifetime that ends by itself is the cheapest control you have on blast radius, and the price of it is that every long-running consumer must know how to renew. That price is paid once, in the client, and not by lengthening everyone's exposure.
- How do you tell an expired credential from a missing permission in the failures?By the shape. An expiry starts at a point in time and refuses everything afterwards, including harmless reads and calls that succeeded a minute earlier. A missing permission is present from the first attempt and refuses one action while the rest keep working. Widening a grant to cure the first shape leaves a permission nobody needed.
- The unit of work cannot finish inside the maximum session lifetime — now what?Make the unit smaller. Checkpoint after each piece, resume by skipping finished pieces, and let each piece run under a freshly acquired session. Where that is genuinely impossible, treat the caller as a named exception with an owner and a review date rather than raising the ceiling for the whole estate.
saying these in an interview costs you the question
- Reads the refusals as a permissions problem and widens the grant
- Raises the session lifetime until the longest job fits
- Copies the credential values at start-up and never refreshes them
- Refreshes only after a call has already been refused
- Assumes a refresh can extend a session past its maximum lifetime