skip to content

Your six-hour exporter now re-reads its credential every hour and still fails mid-run — what did the re-read alone not fix?

level: middleimportance: should knowfreq 40%

answer

  1. two acts, not one
  2. the variable changed, the session did not
  3. guard before the unit, not after the error
  4. checkpoint boundary, restartable unit
  5. one owner of the value, not four copies

basics

~20 s

The sessions. Re-reading updates a variable; connections authenticated with the previous value are untouched and keep whatever standing they had, and a fresh value reaches nothing until the consumer re-establishes them. Re-resolve and reconnect are two separate acts, and the fix needs both.

solid answer

~40 s

A consumer owes two things, not one. **Re-resolving** asks the store for the current value and its new expiry; **re-authenticating** makes the systems it talks to accept that value. The exporter did the first and skipped the second, so every connection it was already holding still stood on the old credential and still failed at whatever moment it was rebuilt. Where the pair belongs is at a **restartable boundary** — after a checkpoint, before the next unit of work, guarded by a test like *would the next unit run past the moment this credential stops being valid*. Putting it only in the error handler means each failure costs the unit that was in flight; putting it inside the innermost write costs a round trip per record and still leaves the session stale.

code

pseudocode · 10 lines
pseudocode
held = store.read("partner-writer")          # { value, expiresAt }

for chunk in remainingChunks():
    # forward-looking guard: would this chunk run past the expiry?
    if now() + chunk.estimatedSeconds > held.expiresAt:
        held = store.read("partner-writer")   # 1. re-resolve
        pool.closeAll()                       # 2. drop sessions on the old value
        pool.openWith(held.value)             # 3. re-authenticate before writing
    write(chunk)                              # replayable whole if it fails
    checkpoint(chunk.id)

go deeper

for a junior

Recall that reading the credential again and using it are separate steps. The connections a job already has were set up with the old value and will not pick up a new one on their own.

for a middle

Explain the pair — re-resolve then re-authenticate — and why only the second changes what an acceptor will do. Be able to place the pair at a checkpoint boundary and say why a forward-looking guard beats an expiry check.

for a senior

Show the operating detail: units sized so a boundary is cheap, a checkpoint that lets a restart resume, one owner of the value rather than copies in several clients, and an error path kept as a backstop you monitor.

for a principal

Make it a contract other teams inherit rather than advice: a shared client that owns the value and rebuilds sessions, plus a written expectation that no unit of work exceeds a stated fraction of the shortest window a consumer is issued.

## Two separate acts The words get used interchangeably and they are not the same thing: - **Re-resolve** — ask the store again, get the current value and the moment it stops being valid, replace the copy in memory. - **Re-authenticate** — present that value to each system the consumer talks to, so that the session it serves stands on the new credential. Only the second one changes what the acceptor will do. The first changes a variable. A consumer that re-resolves on a timer and never reconnects has added a metric and fixed nothing, which is exactly the shape of the exporter in the question. ## Why re-reading alone changes nothing An established session was authorised at the moment it was set up. The acceptor is not consulting your process's memory; it is serving a session it already decided about. So the fresh value sits in a variable that no open connection will ever use, while the connections that matter continue on the standing of a credential that may already be dead. When one of them is rebuilt — and something always rebuilds one — whether it succeeds depends entirely on whether the reconnect path reads the variable or a copy captured at start-up. The symmetry is worth stating plainly, because candidates usually get exactly one half of it: | Act performed | Open sessions | New connections | |---|---|---| | Re-resolve only | unchanged, still on the old credential | succeed, if the reconnect path reads the new value | | Reconnect only | replaced, but with the same stale value | refused once the window has closed | | Re-resolve then reconnect | replaced and standing on the current value | succeed | ## Where the pair belongs in the loop The placement question is the real content of this leaf, and there are four candidate positions: 1. **At start-up only.** The original defect. One window for a run that outlives it. 2. **Inside the innermost write.** Correct but absurd: a round trip to the store per record, and it *still* does not reconnect anything unless you also tear down the session per record. 3. **In the error handler.** A necessary backstop, not a plan. By the time it fires, the unit of work in flight has already failed, and on a multi-part write that can mean discarding meaningful progress. 4. **At a restartable boundary.** The right answer: after a checkpoint, before the next unit begins, guarded by a forward-looking test. The guard is what makes position 4 cheap. Not *has it expired* but **would the next unit of work run past the moment this credential stops being valid**. That test uses the expiry recorded at resolve time and an estimate of the next unit's duration, and it fires at most once per window rather than once per record. ## What makes the boundary cheap The boundary only works if stopping there costs little, which pushes work back into the job's own design: - **Units small enough** that one is worth redoing — minutes, not hours. - **A checkpoint** that records which units are complete, so a restart resumes rather than repeats. - **Units that can be replayed whole** without corrupting what was already written, so a half-finished unit can simply be run again. - **One place in the code that owns the value**, so re-resolving actually reaches every client, rather than three copies captured in three constructors. That last one is where this quietly fails in real systems. A value read once and passed into several clients at construction time cannot be re-resolved at all — the process is holding four copies and updating one. ## Keep the error path, but demote it None of this removes the need to handle a refusal. Something will eventually be refused: a window shorter than you expected, a unit that ran long, an acceptor that dropped a session. The error path should re-resolve, reconnect and replay the unit — but it is the safety net, and its firing is a signal that the forward-looking guard was wrong. If your logs show the error path doing all the work, the guard's estimate of unit duration is the thing to fix. ## The shape to aim for Read it as a contract the consumer owes: *I know when my credential stops being valid; before each unit of work I check whether the unit fits inside that; if it does not, I take a fresh value and rebuild my sessions before I start; and if I am refused anyway, I take a fresh value, rebuild, and replay the unit.* Everything in this leaf is an implementation of those four clauses.

  • Why is the error handler the wrong place to put the re-resolve and reconnect?
    Because by the time it runs, the unit of work in flight has already been refused, and on a long multi-part write that can discard real progress. Keep it as a backstop that re-resolves, reconnects and replays the unit — but drive the normal case from a forward-looking guard, and treat the backstop firing often as evidence the guard's duration estimate is wrong.
  • A process passes its credential into three clients at construction time — what breaks about re-resolving?
    Each client kept its own copy, so re-resolving updates one variable and leaves three clients on the old value. The consumer needs one owner of the credential that clients ask for the current value, or an explicit rebuild of every client when the value changes. Otherwise the fix appears to be in place and is not.
  • The guard needs the next unit's duration — where does that estimate come from?
    From the job's own history: the recent distribution of unit durations, taken at a high percentile rather than an average, plus a margin. Units are deliberately small, so the estimate does not have to be good — it has to be conservative. If a unit regularly runs past its estimate, the units are too big for the boundary to protect anything.

saying these in an interview costs you the question

  • Believes a fresh value automatically applies to connections already open
  • Puts the re-resolve inside the innermost write and calls it fixed
  • Relies on the error handler as the normal path rather than the backstop
  • Keeps several copies of the value in separately constructed clients
  • Checks whether the credential has expired instead of whether the next unit fits