skip to content

A compromised sensor identity must lose every credential it holds — what does that kill write, and how big does the suppression list get?

level: seniorimportance: should knowfreq 38%

answer

  1. two stores, two writes
  2. rows first, then the cut
  3. one entry, not one per token
  4. entries expire with what they suppress
  5. an empty store fails open silently

basics

~20 s

Revoke every refresh row for that subject, then write one suppression entry keyed by subject and a cut instant so already-minted short-lived tokens are rejected. The list stays small: each entry expires with the tokens it suppresses.

solid answer

~50 s

Two writes, in two stores, in that order. First a ranged update over the refresh table — every row for that `subject`, across every family and every device, set to a terminal `revoked` state, which is why the table needs an index on `subject` and not only on the digest. Second, one suppression entry saying *reject any token for this subject issued at or before instant T*. Order matters: write the cut first and a refresh committing in between mints a token later than the cut, which the cut will never catch. Key the entry by subject and cut rather than by `jti`, because you do not hold the identifiers of tokens you never recorded, and one entry covers them all. Sized that way, the list holds one entry per kill for one token lifetime, not one entry per token.

code

pseudocode · 15 lines
pseudocode
def kill_subject(subject):
    # 1. rows first: after this commits, nothing new can be minted
    cursor = 0
    while batch = credentials.page_by_subject(subject, after=cursor, size=500):
        credentials.set_state(batch, REVOKED)     # terminal, so re-running is a no-op
        cursor = batch.last_id
        progress.record(subject, cursor)          # restart resumes here

    # 2. then the cut, covering every token minted before now
    cut = now()
    suppression.put(
        key   = "subject:" + subject,
        value = {"reject_iat_at_or_before": cut},
        ttl   = MAX_ACCESS_TOKEN_LIFETIME,        # expires with what it suppresses
        keep  = LATEST_CUT_WINS)

go deeper

for a junior

Take away the shape: revoking stored credentials stops new tokens being minted, but tokens already handed out are somewhere else and need their own write to stop working.

for a middle

Explain the two stores and the ordering, and why a subject-plus-cut entry covers tokens whose identifiers you never recorded while per-token entries cannot.

for a senior

Demonstrate the operational half: the index the ranged update needs, the cursor that makes the pass restartable, the sizing argument that keeps the list bounded by incidents, and what an empty store does silently.

for a principal

Own the availability question the design forces: consulting the suppression list on every verification buys a shorter tail and buys a dependency, and someone has to decide which of those the fleet can afford.

A sensor comes back from a depot with its flash visibly read. The decision to kill everything for that identity is made elsewhere — by an operator on a device inventory, by a credential-recovery flow, by an automated risk rule. What this leaf owns is what the token side has to write when that event fires. ## Two stores, two writes The credentials for one subject live in two different places and only one of them is a table you can update: 1. **The refresh rows.** A ranged update: every row with that `subject`, across every family and every device it ever enrolled, set to a terminal state. This needs an index on `subject`; the digest index cannot serve it, because you do not have the digests. 2. **The short-lived tokens already minted.** These are not in any table you own — they are in flight, in device memory and in verifier caches. Nothing you write to the refresh table touches them. Suppressing them takes a second write, to whatever store verification consults. ## Order the writes, or you leave a hole Revoke the rows first, then write the cut. If the cut goes first at instant T, a refresh that was already in flight can commit at T plus a few milliseconds and mint a token whose `iat` is later than T — a token the cut will never suppress and whose row is now dead, so nothing else will catch it either. With the rows dead first, no further mint is possible, and a cut placed at *now* covers everything that did get minted. Use an inclusive comparison on the boundary second, because `iat` has one-second granularity and you would rather over-suppress the second you were already killing. ## Keyed by token identifier, or by subject and a cut | keying | entries per subject-wide kill | covers tokens you never recorded | cost | |---|---|---|---| | one entry per `jti` | one per live token | no | a burst on every bulk kill | | subject plus cut instant | one | yes | verification must know the subject before checking | Per-`jti` suppression is the right shape when you are killing one specific token you have in your hand. A subject-wide kill is the other case, and there the cut entry is the one that works, because the set you are trying to suppress includes tokens your service never wrote down. ## Sizing the list This is what makes the structure cheap enough to keep in a shared store that every verifier reads: - Each entry's lifetime is pinned to the tokens it suppresses — it can be discarded once no token issued before the cut can still be within its own validity. - So the peak size is **kills per token lifetime**, not requests per second. With short-lived tokens measured in minutes, even a stolen pallet of nine hundred sensors killed in one pass leaves nine hundred small entries for those few minutes. - Keyed per `jti` instead, the same pallet would write one entry per live token in one burst, and the structure would be bounded by traffic rather than by incidents. ## Idempotent and restartable A bulk kill will be interrupted — a deploy, a timeout, an operator who runs it twice. 1. Every write is **set-to-terminal-state**, never a toggle, so re-running it changes nothing that already landed. 2. The suppression entry is **keyed**, so writing it twice is writing it once; if a second kill arrives later, keep the *later* cut. 3. Drive the pass with a cursor over the `subject` index and record progress, so a restart resumes rather than starting over. Do not let the pass hold one long transaction across nine hundred sensors. ## What you have not revoked Be honest about the tail. Between the cut being written and the last verifier observing it, a stolen short-lived token is still accepted — the delay is replication or cache freshness if verifiers consult the store, and the token's whole remaining `exp` if they do not consult it at all. And if that store fails over empty, suppression stops silently: no verifier errors, the revoked tokens simply become acceptable again until they expire. The refresh rows stay dead throughout, so nothing new can be minted; the tail is bounded by the token lifetime, which is why that lifetime is the real guarantee and the list is only an accelerator.

  • An operator ends one sensor's enrolment from a device list. What does the token side owe that event?
    The same two writes, scoped to that device: revoke its refresh rows so nothing more can be minted, and write a suppression entry covering the short-lived tokens it already holds. The inventory surface and the record that drove the decision belong elsewhere; these two writes are the token-side consequence, and without the second one the device keeps working for minutes.
  • Why not simply shorten every suppression entry's lifetime to one minute to keep the list tiny?
    Because an entry that expires before the tokens it suppresses is worse than none: the tokens come back to life while everyone believes them dead. The entry's lifetime is not a tuning knob — it is pinned to the longest token validity still outstanding when the cut was written.
  • How do you tell a subject-wide kill actually finished on a pallet of nine hundred sensors?
    Record the pass as a job with a cursor and a count, then verify from the table rather than the log: no row for those subjects left in a non-terminal state, and one suppression entry per subject with a cut at or after the revocation time. A re-run is safe, so re-running is the cheapest proof.

saying these in an interview costs you the question

  • Deleting the refresh rows is enough, the sensor is locked out
  • Write one suppression entry per token identifier for the whole fleet
  • Suppression entries should live forever so nothing is missed
  • Write the cut first, then revoke the rows
  • An empty suppression store makes verification fail closed by itself