A queued canal-gate job runs eight hours after the operator's access token expires — what does its job record carry instead?
answer
- the credential is expired by then
- store the decision, not the bearer
- one action, one resource, one rule
- the job carries its own deadline
- late is wrong, not merely late
basics
~20 sThe job record carries the authorization decision rather than the credential: the subject it acts for, one action on one resource, the rule that allowed it, and a deadline after which the job is dropped rather than run late.
solid answer
~40 sDo not put the operator's access token in the job payload: it is past its `exp` before the worker wakes, and it carries everything that operator may do rather than one gate change. Record the **decision** instead — the subject it is taken on behalf of, one action on one resource, the rule that permitted it at 16:00, when it was decided, and the job's own expiry. Keep it as a row in a store you control and put only its identifier on the queue, so the thing crossing twelve hours is not a bearer credential. The expiry is the field people forget: an order approved for the 04:00 window is wrong at 09:15, so a missed window means dropped and reported, never run late.
code
json · 11 lines{
"recordId": "1f3b7c",
"onBehalfOf": "operator:4417",
"action": "gate.open",
"resource": "canal/lateral-3/gate-14",
"parameters": { "openPercent": 40 },
"basis": "rule:duty-operator-lateral-3",
"decidedAt": "2026-09-19T16:02:11Z",
"expiresAt": "2026-09-20T05:30:00Z",
"notifyOnDenial": "team:canal-ops"
}go deeper
Remember the shape of the problem: the credential that authorized the request has expired by the time the job runs, so something else has to cross the gap. Answering "store the decision, not the token" already puts you ahead of most candidates.
Be able to list the record's fields and defend each one: subject, a single action, a single resource, the rule that permitted it, when it was decided, and when the job itself stops being valid. Explain why that is narrower than any token you could have stored.
Show you have operated this. Name the job's own deadline as a business number distinct from a token lifetime and from queue retention, and say exactly what happens — and who is told — when the window is missed rather than met.
The tradeoff is where authority for non-request work lives at all: a narrow decision record per job, against broad worker permissions plus a queue everybody trusts. Argue the blast radius of each, and what it costs teams to keep the first honest over years.
## The gap the job record has to cross A canal-gate controller takes an instruction from a duty operator at 16:00: **open gate 14 to 40% for the 04:00 watering window**. The operator was signed in, the request carried a live access token, and the service that handled the request authorized the change there and then. Nothing opens yet — the work is queued, and a worker picks it up twelve hours later. By then the access token that authorized the request is long past its `exp`, and the operator may have left the district altogether. The queue has bridged twelve hours; the credential was built to bridge twenty minutes. Whatever the worker acts on at 04:00, it is not that token. ## Three reflexes that do not survive contact - **Serialising the operator's access token into the job payload.** It is expired when the worker wakes, so it does not even work; and it authorizes everything that operator may do, not the one gate change that was approved. - **Minting a token long enough to cover the worst case.** The lifetime is now set by the longest backlog you have ever had rather than by the work, and a broad, long-lived credential is being used to solve a scheduling problem. - **Letting the worker act as itself.** A worker running under its own machine identity acts on behalf of *anyone who can put a message on the queue*. Its permissions have to be the union of every job it might ever run, which quietly makes the queue — not the authorization system — the boundary that decides what happens. ## Record the decision, not the credential What crosses the twelve hours is the **authorization decision**: the fact that at 16:00 a particular principal was allowed to do one particular thing. | field | what it pins down | |---|---| | subject | the principal the action is taken on behalf of | | action + resource | one verb, one object — `open`, gate 14, 40% — never a wildcard | | basis | the rule, role or grant that permitted it at 16:00 | | decided-at | when the decision was taken, so staleness is measurable | | expires-at | the job's own deadline, after which it is dropped rather than run | Two properties make this safe where a token is not. It is **narrow**: it authorizes one action on one resource instead of everything the subject can reach. And it is **not a bearer credential** — a copy of it is worth nothing to anyone who is not the worker reading it out of a store you control, which is why the usual shape is a server-side row with only its identifier on the queue. It is also written by the component that made the decision, and never accepted from the enqueuing caller. A record a caller can compose is a request, not a decision. ## The deadline is the field people forget An irrigation order is not merely *approved*; it is approved **for a window**. Executing it at 09:15 is not a late success — it is a different action that nobody approved. Water on a field at the wrong hour is a loss, and the operator who approved it is no longer watching. So the record carries an expiry of its own, and that clock is none of the others in play: 1. the access token's `exp` — minutes to an hour, and already gone; 2. the queue's own retention or redelivery window — an infrastructure number, chosen by whoever sized the queue; 3. the job's expires-at — a **business** deadline, tied to the watering window, and the only one that answers "is this still the right thing to do?". When it passes, the correct behaviour is to drop the job and tell a named owner. "Ran twelve hours late" and "dropped, nobody knew" are both incidents; a record with an expiry and a notification path avoids the first without creating the second. ## What this record is not It is a **capability the worker acts on**, not an entry in an audit trail. An audit record is written so a human can reconstruct afterwards what happened; this row is read by a machine *before* anything happens, and its contents change what the machine does. Different write path, different retention, and a very different consequence for being wrong — which is why collapsing the two into one artefact is a smell even when their field lists look similar. ## And it does not settle the 04:00 question Recording the decision gets the *what* across the gap; it does not establish that the action is still allowed. Entitlements move while a job waits, and a decision taken at 16:00 is a statement about 16:00. What the record buys you is a bounded, specific action that current authority can be re-checked against — which is precisely why it is worth writing narrowly.
- Why not mint a token at enqueue time whose lifetime covers the queue delay?Its lifetime would be set by the worst backlog you have ever seen rather than by the work, and it would authorize everything the subject may do rather than the one gate change. It is also a bearer credential resting in a durable store for hours. A decision record is narrower, is worth nothing to anyone but the worker reading it, and can be re-checked against current authority when the job actually runs.
- The job is retried four times across an hour — what stops the retries from running past the watering window?The record's own expiry, checked on every attempt before anything else happens. Retries inherit the deadline instead of resetting it, so once the window has passed the job is dropped and reported rather than retried. Queue-level retry limits are an infrastructure backstop and a useful one, but they know nothing about what the business window was.
- How do you stop whoever can write to the queue from composing an enqueue record themselves?Let the component that made the decision write it, and put only its identifier on the queue, so the worker reads the record from a store it trusts rather than from the message body. A record a caller can compose is a request, and treating a request as a decision hands the queue the authority that belongs to the authorization system.
A contractor who has to come back tomorrow is not handed your house key and told to keep it. They get a written work order: this one job, this address, void after Friday. The key is authority to do anything for as long as it is held; the order is authority to do one thing until a date. And the gatekeeper still checks the order against today's list of who may be on site.
saying these in an interview costs you the question
- Just serialise the operator's access token into the job payload.
- Mint a token long enough to cover the deepest backlog we have had.
- Give the worker broad permissions; it is internal, so it is trusted.
- If the job is late, run it anyway — the approval was valid when made.
- The enqueue record is just our audit log entry under another name.
- Put the whole action as free text and let the worker interpret it.