Which signal tells you a running job's credential will expire before the job finishes, rather than after it fails?
answer
- neither side can see it alone
- record the expiry at resolve time
- compare validity left against work left
- percentile duration, never the average
- alert on margin, not on expiry
basics
~20 sThe margin: remaining validity against expected remaining work, both published while the job runs. Emit the moment the held credential stops being valid at resolve time, emit remaining runtime at each checkpoint, and alert when the first is smaller than the second.
solid answer
~50 sYou cannot see this failure coming from the store's side — the store knows what it issued but not how long the holder's work will take — and you cannot see it from the job's side alone, because the job knows its progress but not its deadline. The signal is the **comparison**, so both numbers must exist as telemetry. At resolve time the consumer emits the moment its credential stops being valid, tagged with the job and the host; at each checkpoint it emits how much work is left, taken from a high percentile of its own history rather than an average. The alert is `remaining validity < remaining work + margin`. Alerting on the expiry itself is too late to act on, and alerting on authentication refusals fires only once the work is already lost.
code
pseudocode · 13 lines# at resolve time: publish when this credential stops being valid
held = store.read("partner-writer")
emit("credential_valid_seconds",
value = held.expiresAt - now(),
tags = { job: "nightly-export", holder: hostId }) # times only, never the value
# at every checkpoint: publish what is left of the work
emit("job_remaining_seconds",
value = p95RunSeconds - elapsedSeconds,
tags = { job: "nightly-export" })
# alert when the first is smaller than the second
# credential_valid_seconds < job_remaining_seconds + marginSecondsgo deeper
Recall that the warning does not arrive by itself. Somebody has to record when the credential stops being valid and compare it with how much work is left; neither the store nor the job does that on its own.
Explain where each number comes from — the expiry captured at resolve time, remaining work from the job's own progress — and why the comparison, not either number, is the signal.
Demonstrate the operating judgment: a high percentile rather than an average, margin expressed in units of work, re-emission on every re-resolve, and the refusal alert kept as a backstop that tells you the predictive one was miscalibrated.
Ask what the estate can answer. If nobody can list every holder with its expiry and remaining work, the first such incident is unbounded, and making resolve-time telemetry a default of the shared client is the cheapest fix available.
## Why neither party can see it alone This failure is invisible to both halves of the system, and for symmetrical reasons. - **The issuer** knows exactly what it handed out and until when. It does not know that the process that took it is three hours into a six-hour run. It usually does not even know which host took it. - **The consumer** knows its own progress precisely. It does not, by default, retain the moment its credential stops being valid — that field is read once at resolve time and thrown away with the rest of the response. So the predictive signal does not exist anywhere until somebody creates it, and creating it is one line at resolve time. That is the whole intervention, and it is why teams who have been bitten twice still cannot answer 'which running consumers will expire before they finish'. ## The two numbers **Remaining validity.** At resolve time, record the moment the credential stops being valid and emit the difference from now, tagged with the job name and the holder. Emit the *times*, never the value — this is telemetry, and a credential does not belong in it. Re-emit on every re-resolve so the series steps up rather than running to zero and disappearing. **Remaining work.** At each checkpoint, emit an estimate of what is left. Where the job has a countable unit — chunks, partitions, files — the estimate is the count remaining times a recent per-unit duration. Where it does not, use the job's own historical run length minus elapsed time. Take the percentile, not the mean. A job that averages four hours against a five-hour window looks safe and is not: the average night is not the night that fails. If the slowest one run in twenty takes five and a half hours, the margin is already negative one night a month. ## Alert on the margin, not on the expiry | Alert on | Fires | Worth having? | |---|---|---| | Remaining validity < remaining work + margin | while there is still time to act | yes — this is the signal | | The credential's window closing | at the deadline, with no lead time | no — nothing can be done with it | | Authentication refusals from the acceptor | after the unit of work has failed | yes, as a backstop and for the post-mortem | | The job failing | after hours of work are lost | that is the incident, not the warning | The refusal alert is worth keeping — it does fire, and it is how you learn the predictive one was miscalibrated. It is simply not the signal this leaf is about, because by then the answer to 'what do we do' is 'restart and pay for the run again'. ## What the estate-level view looks like Once both series exist, the useful query is a list, not a graph: **every consumer currently holding a credential, the moment each one stops being valid, and the work each has left.** Sorted by margin, that list is a queue of tonight's incidents in the order they will happen. It also answers the question that always comes up after the first one of these — *how many other jobs are in this position* — which before the telemetry existed was answerable only by reading code. ## Traps in practice - **Progress reported in rows rather than time.** '80% of rows written' says nothing about the remaining 20% if the last partition is the big one. Convert to time using recent per-unit durations. - **A margin that is a fixed number of minutes.** Ten minutes of margin means nothing to a job whose units take twenty. Express it in units of work, or as a multiple of the longest recent unit. - **A job that never re-emits.** If the series is written once at start-up, it decays to nothing useful the moment the job re-resolves. Emit on every resolve. - **Blaming the estimate for the failure.** The estimate being wrong is a tuning problem. Having no estimate at all is the defect. ## Where this stops The signal tells you the collision is coming. It does not prevent it — a consumer that cannot re-resolve and reconnect will still fail, just with warning. The value of the warning is that the choice becomes yours: re-run the job now while it is cheap, shorten tonight's workload, or stand by. That is a much better position than reading the failure out of the job's own error at dawn.
- Why is an alert on the credential's expiry itself close to useless here?Because it carries no lead time. It fires at the deadline, when the only options left are to let the job fail or kill it, and it fires for every holder whether or not the holder still has work. The margin alert fires while there is time to re-run cheaply, shorten the night's workload, or simply stand by informed.
- The job averages four hours against a five-hour window — why is that margin unsafe?Because the average night is not the night that fails. If one run in twenty takes five and a half hours, the margin is negative roughly once a month, and the alert built on the mean will never have fired. Size the estimate at a high percentile of recent runs and add margin in units of work, not in fixed minutes.
- What does the estate-level version of this give you that a per-job alert does not?A list of every holder, when each one's credential stops being valid, and the work each has left — sorted by margin. That is the queue of the coming week's incidents, and it answers 'how many other jobs are in this position', which otherwise requires reading every consumer's source.
saying these in an interview costs you the question
- Expects the store to warn holders whose work will outlast the credential
- Alerts on the expiry moment and calls that early warning
- Treats authentication refusals as the predictive signal
- Sizes the margin from the average run rather than a slow one
- Emits the credential itself alongside the timing metrics
- Reports progress in rows and converts nothing to time