skip to content

Your estate runs batch jobs longer than the validity window of the credentials they hold — what standard do you set for them?

level: principalimportance: should knowfreq 30%

answer

  1. constrain the unit, not the job
  2. restartable boundaries make length irrelevant
  3. a shared client cannot reach every consumer
  4. a longer window is a bridge, not a fix
  5. count the exemptions each month

basics

~20 s

Make the unit of work, not the job, the thing that must fit inside a window: checkpointed units shorter than the shortest window issued, with a re-resolve and reconnect between them, delivered in a shared client so teams inherit it — and name the consumers that client cannot reach.

solid answer

~50 s

There are three moves and they are not equivalent. **Shorten the unit of work** so no unit outlives a window, and re-resolve and reconnect at each boundary — this is the only one that scales to jobs of arbitrary length, and it costs the engineering to make units restartable. **Put the pair in a shared client** so that consumers inherit the behaviour rather than each team building it; this costs you ownership of the client and a rollout, and it cannot reach a consumer that hands the value to a separate process or a third-party client. **Ask for a longer window** for that class of consumer — that is a change to the issuance contract, negotiated with whoever sets it, and it buys the margin with a wider exposure rather than removing the failure. Write the standard as a rule about units, publish which consumers are exempt, and measure the exemptions.

go deeper

for a junior

Recall the reframing: the fix is not a longer-lived credential but a smaller unit of work, so that a job of any length is a sequence of units each short enough to finish inside one window.

for a middle

Explain why unit size is the controllable variable and total runtime is not, and what a boundary needs to be cheap — a checkpoint, a replayable unit, and one owner of the credential inside the process.

for a senior

Show the operating side: which consumers a shared client cannot reach, how the rule is stated so compliance is measurable, and why the error path stays as a backstop rather than becoming the mechanism.

for a principal

Own the tradeoff explicitly — a job that stumbles cheaply against one that dies expensively, a uniform rule against per-class rules — and commit to the measurement, because an exemption register that grows means the standard was routed around.

## Restate the problem so it has a solution 'Jobs that run longer than their credentials' has no fix, because you cannot make a six-hour job shorter by policy. The version that does have a fix is: **no single unit of work may outlive the credential it holds.** That reframing is most of the value of the standard, because it moves the requirement from the job's total runtime — which is set by data volume and which you do not control — to the unit size, which is a design choice. A job of any length then becomes legal: it is a sequence of units, each comfortably inside a window, with a re-resolve and a reconnect at each boundary. ## The three moves, and what each costs | Move | What it actually fixes | What it costs | Where it stops | |---|---|---|---| | Shorten the unit; re-resolve and reconnect at the boundary | any run length, permanently | making units restartable and checkpointed — real engineering in each job | jobs with a genuinely indivisible unit | | Put the pair in a shared client every consumer uses | most consumers, without per-team work | you own the client, the rollout and the migration | a consumer that hands the value to a separate process or a client you do not control | | Ask for a longer window for this class of consumer | tonight, and buys time | a wider exposure, and a change to the issuance contract owned by whoever sets it | it postpones rather than removes the failure | The third belongs in the plan as a **bridge**, honestly labelled. It is not yours to grant, it trades a security property for an operational one, and a standard that quietly leans on it has chosen a longer-lived credential across the estate without saying so. ## What the shared client can and cannot reach This is where most versions of this standard are over-claimed. A shared client fixes a consumer that asks it for the value and lets it own the sessions. It does not fix: - A consumer that reads the value and **passes it to a separate process** it starts — a child process, a sidecar task, an external tool invoked once per run. - A consumer that hands the value to a **third-party client library** which captures it at construction and offers no way to replace it. - A consumer that **renders the value into a file or an environment** at start-up and never looks again. - Anything whose session is **established by a component you do not own**. Name those categories in the standard and require them to be declared, because they are exactly the population that will produce the next incident. Whatever their count is, it is the honest measure of how far the shared client actually reaches. ## Writing the rule so it can be checked A standard nobody can evaluate is advice. Make each clause checkable: 1. **A unit of work must be no longer than a stated fraction of the shortest window a consumer of this class can be issued.** State the fraction and state the shortest window you assumed — if the shortest is one hour and the fraction is a quarter, units run fifteen minutes. 2. **Each consumer must emit, at resolve time, the moment its credential stops being valid**, and, at each checkpoint, its remaining work. 3. **Each consumer must re-resolve and re-authenticate at a unit boundary**, with the error path as a backstop rather than the mechanism. 4. **A consumer that cannot comply must declare itself**, with the reason and an owner. Clause 4 is the one that makes the standard survive contact. There will be non-compliant consumers; a register of them is worth more than a rule they silently ignore. ## The judgment call you are actually making Underneath the mechanics is a choice about which failure you prefer. Small units mean frequent, cheap boundaries and the occasional restarted unit — a job that stumbles. Large units mean a smooth run that, when it collides with a window, loses hours of work — a job that dies. The first is worse on any given quiet night and enormously better across a year. The second judgment is whether the rule is uniform. A single fraction across the estate is easy to state and will be wrong for the extremes: it over-constrains a job whose natural unit is seconds and under-serves one whose indivisible unit is an hour. Per-class rules are more honest and cost you the classification. Either is defensible; choosing without saying which you chose is not. ## How you know it worked Count three things, monthly: consumers whose longest observed unit exceeds the rule; consumers emitting no resolve-time expiry at all; and declared exemptions. A standard whose exemption list is growing has not been adopted — it has been routed around.

  • Why not simply ask for credentials that outlive the longest job?
    Because it buys the margin with exposure and it does not scale: the longest job grows, and the window then has to grow with it forever. It is also not your decision — the window is set by whoever owns the issuance contract. Use it as a declared bridge while units are made restartable, not as the standard.
  • Which consumers will a shared client's re-resolve-and-reconnect logic never fix?
    Those that pass the value out of the client's reach: a child process started with it, a third-party library that captures it at construction, a rendered file or environment read once at start-up, or a session opened by a component you do not own. Require these to declare themselves; their count is how far the client really reaches.
  • How do you set the fraction of the window a unit of work may occupy?
    Work from the shortest window a consumer of that class can be issued, not the typical one, and from a high percentile of unit duration rather than the mean. State both assumptions in the rule. A quarter of the shortest window is a defensible starting point because it tolerates a unit running several times longer than expected before anything is at risk.
  • What does a growing exemption register tell you?
    That the standard is being routed around rather than adopted, and usually that the unit-size rule is unreachable for a class of jobs whose natural unit is genuinely large. Treat it as a design signal: either that class needs a different rule, or it needs engineering investment to become divisible.

saying these in an interview costs you the question

  • Presents a longer validity window as the standard rather than a declared bridge
  • Claims a shared client covers every consumer in the estate
  • Writes the rule against total job runtime instead of unit size
  • Sizes the rule from the typical window rather than the shortest issued
  • Leaves no register of consumers that cannot comply
  • Relies on automatic whole-job restarts as the compliance story