skip to content

What revocation-lag number do you sign for on a 40,000-sensor fleet, and which part of it can this service actually shorten?

level: principalimportance: should knowfreq 30%

answer

  1. revoked is a distribution, not an instant
  2. two promises hide in one word
  3. the dominant term is the token lifetime
  4. freshness of the consulted store
  5. prove it with a canary revocation

basics

~20 s

Sign the worst case, not the write. Revocation lag is a minted token's remaining lifetime plus the time a suppression entry needs to reach every verifier; only the second term belongs to this service, and shortening it buys a dependency.

solid answer

~50 s

The number an operator cares about is *how long after we press revoke does a stolen credential stop working*, and it decomposes into three terms. The refresh rows die at commit, so nothing new is minted — that term is effectively zero. A token already minted survives until its own `exp`, and that lifetime is set on the issuing side, not here; on a battery fleet it is held up by radio wakeups, so it will not simply be dialled down for you. What this service owns is the third term: whether verification consults a suppression entry at all, and how fresh that consultation is. Buy the shorter tail and you have made verification depend on a shared store. Then prove the number with a canary revocation measured at a verifier, rather than asserting it.

code

pseudocode · 13 lines
pseudocode
# canary: does the signed budget still hold?
identity   = enrol_throwaway_sensor()
token      = identity.refresh()             # a live short-lived token
t0         = now()
kill_subject(identity.subject)              # rows, then the cut

while now() - t0 < BUDGET * 2:
    if verifier.verify(token) == REJECTED:  # measured at a verifier, not at the writer
        record("revocation_lag_seconds", now() - t0)
        return
    wait(250.ms)

alert("revocation lag exceeded twice the signed budget")

go deeper

for a junior

Take away that revoking stored credentials stops new ones being issued, while a token already handed out keeps working for a while. Ask how long that while is before believing a revoke button.

for a middle

Be able to name the three terms and say which one dominates, and explain why a cached suppression answer sets the real freshness of the whole budget.

for a senior

Show how you would measure it: a canary revocation timed at a verifier, alarms on the p99, and a re-measure after any change to lifetime, cache or topology.

for a principal

Own the commitment itself — who signs the number, what fail-open versus fail-closed costs a fleet that cannot be visited, and which lever lives with the issuing side rather than here.

"Revoked" sounds like an instant. On forty thousand sensors it is a distribution with a tail, and a business that has signed a cold-chain commitment will eventually ask what the tail is. The useful exercise is to write the number down, decompose it, and say which term anyone can actually move. ## What "revoked" actually promises Two different promises hide in the word, and only one of them is instant: - **No new authority.** Once the refresh rows are in a terminal state, that subject cannot mint anything again. This is a committed write, so it holds from the moment the transaction lands. - **No existing authority.** A short-lived token already in a device's memory keeps being presented and keeps being accepted, because verification is a signature check over a self-contained token. Stopping it early takes a suppression entry that the verifier actually reads. Any statement of the form "we revoked it" that does not say which of the two it means is not a budget. ## The three terms | term | typical size | who owns the lever | what shortening it costs | |---|---|---|---| | stop minting | milliseconds | this service | nothing; it is a committed write | | a minted token's remaining validity | up to its full lifetime | the issuing side | shorter lifetimes mean more radio wakeups per sensor per day | | suppression reaching every verifier | freshness of the consulted store | this service | verification gains a dependency on that store | The honest headline is that the second term dominates, and it is not this leaf's lever. A fleet that wakes to trade credentials is spending its battery on radio-on time; cutting the token lifetime multiplies those wakes, and on a poor link each one costs several attempts. That trade is made where tokens are minted, and the right move here is to take the lifetime as given and state the budget against it. ## The lever this service does own What is genuinely yours is the shape of the third term: 1. **Do verifiers consult a suppression entry at all?** If not, the budget is simply the token lifetime, and you should say so out loud rather than implying a revoke button does more than it does. 2. **How fresh is the consultation?** A synchronous read gives you replication lag; a cached answer gives you the cache's freshness window, which is the real number even though nobody wrote it down. 3. **What happens when that store is unreachable?** Failing closed protects the budget and can strand forty thousand sensors on a bad afternoon; failing open keeps the fleet reporting and silently suspends the budget. Pick one deliberately and write which one you picked next to the number. The refresh credential's own lifetime is also yours, and it trades differently: a long one means a sensor asleep in a container hold for six weeks can still come back and refresh, while a short one forces re-enrolment — a provisioning round-trip or, worse, a field visit — for anything that misses its window. It does not extend an attacker's window, because the row is dead server-side either way. ## Measuring that the number holds A budget nobody measures decays silently, usually the first time somebody changes a cache setting: - Run a **canary**: periodically revoke a throwaway identity that holds a live token, then measure at a verifier how long until the first rejection. Alarm on the p99 of that measurement against the signed number. - Track the fraction of verifications that reached the suppression store versus answered from a cached value, because that ratio *is* the freshness term. - Re-run the canary after any change to token lifetime, cache freshness or store topology; those are the three inputs, and all three are edited by people who are not thinking about revocation. ## What you cannot fix, and what you do instead You cannot reach a token that is already in an attacker's hands, and you cannot tell a sensor asleep in a freezer anything at all. What you can do is bound the window, publish the bound, and make the bound observable. One more failure mode belongs in the same conversation: restoring the credential table from a backup taken before a revocation silently returns those rows to active, so any restore has to be followed by replaying the revocations made since the snapshot. That is the kind of detail a signed number forces you to discover before an auditor does.

  • Verification currently consults no suppression store. What is the revocation-lag number then?
    Exactly the token lifetime, worst case, every time. That is a defensible budget if the lifetime is short and the risk accepted, but it must be stated as the number rather than implied away by the existence of a revoke action. Anyone who believes revoke is instant will size an incident response wrongly.
  • Why does restoring the refresh-credential table from a backup threaten the budget?
    Because state is restored as of the snapshot: rows revoked after it come back active and can mint again. The restore procedure has to replay every revocation made since the snapshot, which means those events need to be durable somewhere other than the table being restored.
  • Who should sign the number, and what changes when nobody does?
    Whoever owns the incident-response commitment, because the number is a promise to them, not a property of the code. Without a signature it is an emergent value: a cache freshness change or a token-lifetime tweak doubles it quietly, and the first time anyone measures it is during the incident it was written for.

saying these in an interview costs you the question

  • Revocation is instant once the row is updated
  • Shortening the token lifetime is free, just set it lower
  • Caching the suppression answer does not affect the budget
  • A restored credential table is safe because the rows were revoked
  • Fail-open on the suppression store is not a decision anyone makes